The problem this solves
Support cost centers don't just serve the business β they often serve each other. IT runs the servers Facilities' building-management system depends on; Facilities provides the office space IT's own staff sit in. Allocating IT's cost to stores while ignoring what it owes Facilities (and vice versa) systematically understates both centers' true fully-loaded cost β and worse, which method you use to handle this circularity changes how much of the total ends up "labeled" IT versus Facilities by a material amount, even though the total dollars reaching stores never changes.
This matters whenever cost centers get compared, budgeted, or benchmarked against each other: a support function that looks cheap because a reciprocal relationship went unaccounted for isn't actually cheap β the cost is real, just mislabeled.
Audience
Cost accounting and FP&A teams who own the allocation methodology for internal service/support functions. IT and Facilities finance business partners who need to understand why their fully-loaded cost differs from their direct budget. Anyone deciding which allocation method (Direct, Step-Down, or Reciprocal) their organization should standardize on.
Two cost centers, a mutual service relationship
| Center | Direct Cost | Serves the Other | Serves Stores |
|---|---|---|---|
| NLQ query: "Show the reciprocal method allocation" or "Direct method for IT and Facilities" β DeepSeek resolves the method plus year/scenario | |||
| IT & Digital Services | 1.5% of total revenue | 8% (shared data centers, in Facilities-managed buildings) | 92% |
| Facilities & Real Estate | 2.0% of total revenue | 12% (IT department's own office space) | 88% |
Both centers' direct-cost bases are sized as a percentage of total company revenue at runtime, matching the same pattern the Allocation Waterfall pools use.
The same total, four different labelings
- KPI row: the selected method's ITβStores, FacilitiesβStores, and combined total
- Stacked bar chart: all four methods compared side by side β visually confirms the total bar height never changes, only the IT/Facilities split within it
- Worked math panel: the actual equations for whichever method is selected, filled in with real numbers β including the simultaneous-equation solve for Reciprocal
- Comparison table: all four methods' numbers in one place, selected method highlighted
The harder case the Allocation Waterfall doesn't cover
Allocation Waterfall (UC13) allocates from cost pools down to entities β a one-directional cascade. This use case handles the case where two cost centers allocate to each other as well as downstream, which the simple waterfall model can't represent. Both feed into the same underlying idea: cost has to land somewhere defensible, and the method used to get it there has to be explainable, not just computable.
Recommended integration points
- Standardizing allocation methodology: before an organization picks Direct, Step-Down, or Reciprocal as its house standard, this shows exactly what's traded off β Direct is simplest but least accurate; Reciprocal is most accurate but requires solving a system every period
- Cross-charge disputes between support functions: when IT and Facilities disagree about whose budget should absorb a shared cost, the Reciprocal method's simultaneous-equation solve is the defensible answer, not a negotiated split
- Benchmarking support function cost: comparing IT's cost-per-store against an industry benchmark only makes sense if the allocation method matches what the benchmark assumes
The numbers
The honest checklist
- βTwo or more of your support/service cost centers genuinely serve each other, not just downstream business units
- βYou need to defend a cost-center's fully-loaded cost figure to someone who will ask "why this number and not a different one"
- βYou're choosing (or reconsidering) which allocation method to standardize on and want to see the actual dollar impact of that choice, not just the theory
- βYour cost centers only allocate downstream to final entities, with no service relationship between the centers themselves β that's the simpler case in UC13
- βYou need more than two reciprocal centers modeled simultaneously β this demo's 2Γ2 system illustrates the mechanics; production systems with 3+ mutually-reciprocal centers need a full matrix solve, not the closed-form algebra shown here
Try it now
The live demo recomputes all four methods in real time and highlights the one currently selected.
Things to try
- Switch between all four methods and watch the stacked bar chart β the total bar height never moves, only where the IT/Facilities split line falls within it
- Type "Solve the simultaneous equations for reciprocal costing" and read the worked-math panel β check that
IT Γ 0.92 + FAC Γ 0.88equals the same total every other method produces
The circularity is the whole problem
A one-directional allocation (pool β entities) is a closed-form calculation with no ambiguity. A reciprocal relationship β IT's true cost depends on Facilities' true cost, which depends on IT's true cost β isn't closed-form unless you either break the circularity by fiat (Direct, Step-Down) or solve it as a system (Reciprocal). This demo exists specifically to make that choice, and its consequences, visible and comparable rather than abstract.
Method branch β four independent calculations
- Returns itDirect, facDirect — each a % of total revenue
- Branches on method — see below
- facToStores = facDirect
- Nothing flows back to IT once closed
- Nothing flows back to Facilities once closed
- FAC = facDirect + pIT2Fac·IT (solved simultaneously)
- KPI row + stacked bar (all 4 methods)
- Worked-math panel
- Comparison table
App UI β Component breakdown
| Component | Behaviour |
|---|---|
| Method selector | Direct, Step-Down (IT first), Step-Down (Facilities first), Reciprocal β defaults to Reciprocal, the most defensible method |
| KPI row | Selected method's ITβStores, FacilitiesβStores, and combined total |
| Stacked bar chart | All 4 methods always shown together, so the "same total, different split" property is visible without needing to remember the previous method's numbers |
| Worked-math panel | Renders the actual equations for the selected method with real computed numbers substituted in β not a generic formula, the live result |
| Comparison table | All 4 rows, selected method's row highlighted |
What actually differs
| Method | How it handles the circularity | Accuracy |
|---|---|---|
| Direct | Ignores it entirely β each center's direct cost goes 100% to stores as if the reciprocal relationship didn't exist | Lowest β systematically mislabels cost that actually flowed through the other center |
| Step-Down (IT first) | IT allocates out (to Facilities and stores) and is then "closed" β nothing flows back into it | Better than Direct, but order-dependent: the center processed first is understated |
| Step-Down (Facilities first) | Same idea, reversed order | Same limitation, opposite bias |
| Reciprocal | Solves the 2Γ2 simultaneous system so both centers' fully-loaded cost already includes what they received back from each other | Highest β no sequencing artifact, no ignored relationship |
Solving the 2Γ2 system in closed form
IT = itDirect + pFac2IT Γ FAC
FAC = facDirect + pIT2Fac Γ IT
// substitute FAC into the IT equation:
IT = itDirect + pFac2IT Γ (facDirect + pIT2Fac Γ IT)
IT Γ (1 - pFac2IT Γ pIT2Fac) = itDirect + pFac2IT Γ facDirect
IT = (itDirect + pFac2IT Γ facDirect) / (1 - pFac2IT Γ pIT2Fac)
FAC = facDirect + pIT2Fac Γ IT // back-substitute
With pFac2IT = 0.12 and pIT2Fac = 0.08, the cross-term pFac2IT Γ pIT2Fac = 0.0096 is small, so the closed-form solve converges to essentially the same answer an iterative "allocate back and forth until it stabilizes" approach would reach β but the closed-form version is exact and instant, with no convergence tolerance to tune. A larger, N-center reciprocal system would need a full matrix solve (Gauss-Jordan or similar) instead of this 2-variable algebra; this demo intentionally keeps to 2 centers so the math stays inspectable in a worked-math panel rather than a matrix inversion nobody can eyeball.
Redistribution, not creation or destruction
Every method starts from the same two direct-cost numbers (itDirect + facDirect) and every method's store-bound total works out to exactly that same sum β verified directly in testing across all four methods for the same year/scenario. This is the single most important sanity check on the whole demo: if a method's total-to-stores ever diverged from the others, that would mean cost was being invented or lost, not just relabeled, which would be a bug, not a legitimate methodology difference. The four methods differ only in which center gets credited with how much of that fixed total β never in the total itself.
A method-selection schema, not an entity-lookup one
Unlike most of the site's other use cases, Reciprocal's NLQ schema β {intent, method, year, scenario, confidence} β has no entity field at all: there's nothing to look up, only a method to select from a fixed enum of 4 values. The few-shot examples and the semantic-fallback keyword matcher both key off phrases like "direct," "step down...IT first," and "reciprocal/simultaneous" rather than any lookup table, since the entire space of valid answers is 4 strings.
Cost model β At-scale projections
LLM cost is flat per query at roughly $0.0002 β the smallest lookup space of any use case on the site (4 fixed method strings). The reciprocal solve itself is O(1) β a closed-form 2-variable algebraic solution, not an iteration β so it stays instant regardless of company size. The real production cost driver would be extending this to N>2 reciprocal centers, which requires a genuine matrix solve rather than this demo's closed-form shortcut.
Key files
| File | Role |
|---|---|
epm-nlq-src/assets/epm-pcmcs-data.js | RECIPROCAL_CENTERS, RECIPROCAL_METHODS, getReciprocalDirectCosts(), computeReciprocalAllocation() |
epm-nlq-src/pages/reciprocal.html | UI: NLQ bar, method/year/scenario filters, KPI row, stacked chart, worked-math panel, comparison table |
epm-nlq-src/assets/epm-data.js | Shared genIS() β both centers' direct cost bases are sized off total company revenue |
functions/api/nlq-query.js | Shared NLQ endpoint; Reciprocal uses useCase: "reciprocal" with a method-enum schema |
Tech stack β Every tool in this build
| Layer | Tool | Why |
|---|---|---|
| Data layer | epm-pcmcs-data.js (vanilla JS) | Closed-form algebraic solve β the LLM only resolves the method enum |
| LLM | DeepSeek V3 | Shared NLQ endpoint; smallest schema of the twelve use cases β a 4-value enum plus year/scenario |
| Charts | Chart.js 4.x | Stacked bar comparing all 4 methods, shared library with Capital UC6, Ownership UC9 |
| Edge hosting | Cloudflare Pages | Static file, no server compute needed for this use case |
| Build | Eleventy v3.1.5 | Copies epm-nlq-src/pages/ and epm-nlq-src/assets/ to _site/ verbatim |
Known attack surfaces
| Threat | Mitigation in this build |
|---|---|
| Grand total silently diverging between methods (cost creation/destruction) | Verified directly in testing: all four methods' totalToStores match to within $0.01 for the same year/scenario |
| Reciprocal solve diverging or producing a negative cost | With pFac2IT Γ pIT2Fac = 0.0096, the denominator (1 - 0.0096) stays safely positive and close to 1 for any realistic percentage pair in this model β the closed-form solve is well-conditioned by construction |
| Step-down method silently applied in the wrong order | stepdown_it and stepdown_fac are two explicit, separately-coded branches β there's no shared "step-down" function with an order parameter that could be passed backwards |
Guardrails β What prevents bad reciprocal math
- Method is an explicit enum, never inferred: exactly 4 valid
methodIdvalues, validated againstSCHEMA.reciprocalMethodsbefore the calculation runs β an unrecognized method string falls through to the safe default (reciprocal) rather than silently mis-branching - Every method's output is cross-checked against the same total: the demo computes all four methods every time it renders (not just the selected one), so a regression that broke the total-preservation property would be immediately visible in the comparison table, not hidden behind a single-method view
- pctToOther + pctToStores is enforced to sum to 1.0 for each center in
RECIPROCAL_CENTERSβ a center that "serves itself" more or less than 100% of its output would silently create or destroy cost
Who is asking, and what are they allowed to see?
The demo answers neither question — it has a cookie gate and no notion of a user. In production these are the two questions everything else rests on, and they have different answers: authentication is who you are, authorization is what you may see. Corporate SSO settles the first. Only Oracle EPM Cloud can settle the second, and the single most important rule in this section is that this application must never become the place where that decision is made.
9.1 · The identity chain, end to end
- MFA and Conditional Access are enforced here — device compliance, location, risk signals
- Returns an ID token (who the user is) and an access token (what they may call)
- Group membership arrives as a claim; the application never handles a password
- Verify signature, issuer, audience and expiry against the IdP’s published keys
- Read the group claims — there is no local user table and no local role table
- Short-lived access token with refresh-token rotation; session timeout set to the data classification
- The API is called as the user
- EPM enforces its own security natively
- Audit trail names the real user
- Preferred where the API supports it
- The application becomes the enforcement point
- Entitlements fetched separately, applied in one audited place
- Simpler and cacheable — and a filtering bug is a data breach
- Roles: Service Administrator, Power User, User, Viewer — assigned to groups, never to individuals
- Data level: PCMCS application and POV access, with rule-set maintenance restricted to the cost-accounting owner
- Group → role mapping lives in the platform, not in this application
- The user sees exactly what they would see logging into the source system directly — no more
- Every query logged against the real end user, never a shared account
9.2 · Federating the corporate identity provider
Oracle EPM Cloud does not replace your directory — it trusts it. The EPM Cloud identity domain is federated with the corporate IdP so authentication happens where it already happens, under policies security has already written.
| Identity provider | Protocol | Notes |
|---|---|---|
| Microsoft Entra ID (formerly Azure AD) | SAML 2.0 or OIDC | The common case. Conditional Access, MFA and device compliance are enforced at Entra and inherited automatically. On-premises Active Directory federates through Entra Connect rather than being integrated directly. |
| Okta | SAML 2.0 or OIDC | Same pattern; Okta groups drive EPM roles through SCIM provisioning. |
| OCI IAM (identity domains) | Native | Already present with Oracle EPM Cloud. Can be the primary IdP for a smaller estate, or a federated spoke of Entra/Okta for a larger one. |
| AD FS | SAML 2.0 | Still seen where the estate is not yet cloud-first. Works, but you inherit the on-premises availability of the token service — if AD FS is down, nobody logs in. |
For the browser application itself, use OIDC Authorization Code flow with PKCE. Not the implicit flow, which is deprecated and leaks tokens through the URL, and never a resource-owner password grant — a finance tool should not be capable of handling a password at all.
9.3 · From group membership to EPM roles
Roles are granted to groups, never to individuals, and the groups come from the directory. That one discipline is what makes joiner/mover/leaver work without anyone having to remember this application exists.
Entra ID group β EPM role / entitlement
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
FIN-EPM-Analysts β Planning User
FIN-EPM-Controllers-EMEA β Power User + EMEA data scope
FIN-EPM-Admins β Service Administrator
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Provisioned by SCIM. Remove the user from the group and the
entitlement disappears on the next sync β including here.
For this use case the relevant native entitlement is: PCMCS User; changing the costing method is a model change under change control, not a user action.
9.4 · The architectural decision: who enforces?
This is the choice that determines whether the deployment is defensible. Both patterns appear in the diagram above; the difference is where the security boundary actually sits.
| Pattern A — identity propagation | Pattern B — service account + filtering | |
|---|---|---|
| How | The user’s token is exchanged (OAuth 2.0 on-behalf-of) for one scoped to the EPM API; calls are made as the user | A single read-only integration account calls the API; the application filters the results |
| Enforcement point | Oracle EPM Cloud | This application |
| Audit trail shows | The real end user | The service account — you must log the real user separately |
| Failure mode | Token plumbing is more complex; per-user rate limits apply | A filtering bug is a data breach, and the entitlement copy drifts from reality |
| Verdict | Prefer this wherever the API supports user-token authentication | Acceptable with discipline: narrowest possible service account, filtering centralised in one tested place, real user in every log line |
The shortcut to refuse. Pattern B built with a Service Administrator account and no filtering at all is the most common way this gets delivered, because it works perfectly in UAT — testers are usually over-entitled, so nobody notices that everyone can see everything. It fails at the first access review, and by then it is in production with real users depending on it.
9.5 · Data-level security is the part that matters
Role membership decides whether a user can open the application. It does not decide which rows they get back, and confusing the two is the most expensive mistake available here.
- For this use case: PCMCS application and POV access, with rule-set maintenance restricted to the cost-accounting owner.
- Apply it before aggregation, not after. Filtering a total that has already been computed across entities the user cannot see still leaks the total.
- The NLQ layer needs its own check. Layer 4 already validates that the resolved point of view uses approved members; production adds a second test — that the resolved POV sits inside this user’s scope — and it runs before the data call, not after. A natural-language interface is very good at asking for things politely; the authorization check must not care how the question was phrased.
- Fail closed. If entitlements cannot be resolved, return nothing and say so. An empty result is a support ticket; a permissive default is an incident.
9.6 · Provisioning, sessions and the leaver problem
- SCIM provisioning from Entra or Okta into the EPM Cloud identity domain (OCI IAM, formerly IDCS), covering joiner, mover and leaver. The mover is the case people forget — somebody changing region should lose the old scope, not accumulate both.
- No local user store. If this application keeps its own copy of who may do what, a leaver keeps access until somebody remembers to update it. Nobody ever does.
- Short-lived access tokens with refresh-token rotation; align session timeout with the data classification rather than with convenience.
- MFA and Conditional Access at the IdP — not reimplemented here. Device compliance and location policy come free with federation.
- Quarterly recertification of both the groups that grant access and the service account’s own entitlements, evidenced and signed.
- Break-glass access is a named, monitored, time-boxed account — never a shared credential in a password manager.
9.7 · What this means for Reciprocal Costing
| Concern | Answer for this use case |
|---|---|
| Native entitlement required | PCMCS User; changing the costing method is a model change under change control, not a user action |
| Data-level control | PCMCS application and POV access, with rule-set maintenance restricted to the cost-accounting owner |
| Use-case-specific sensitivity | Method comparison is politically loaded — the choice moves hundreds of thousands between cost centres. Keep the side-by-side view for policy review and restrict it accordingly; published cost-centre numbers must come from the single approved method, not from whichever method a user happened to select. |
9.8 · Security configuration checklist
- ✓Oracle EPM Cloud federated with the corporate IdP over SAML 2.0 or OIDC; the cookie gate removed entirely
- ✓Browser app uses OIDC Authorization Code + PKCE — no implicit flow, no password grant
- ✓MFA and Conditional Access enforced at the IdP, not reimplemented in the application
- ✓Roles granted to directory groups, never to individuals; SCIM covers joiner, mover and leaver
- ✓Enforcement pattern chosen deliberately — Pattern A where the API supports it, or Pattern B with filtering centralised and tested
- ✓Data-level security applied before aggregation, and the resolved POV checked against the user’s scope before the data call
- ✓No local user table and no local role table anywhere in the application
- ✓Every query logged against the real end user, even when a service account makes the call
- ✓Authorization failures fail closed and are logged as security events rather than swallowed
- ✓Quarterly recertification of access groups and of the service account’s own entitlements
From demo to a governed enterprise deployment
Everything above runs on synthetic data, a public LLM API key, a cookie gate, and no audit trail — deliberately, so the mechanics are inspectable. Taking Reciprocal Costing to production is not a rewrite; the 4-layer pipeline and the data-layer contract survive intact. It is a controlled-change program across six workstreams: architecture, LLM platform, security, SOX/audit, environment promotion, and operations. This section is the checklist we run with clients.
10.1 · Production reference architecture
- Terminates SSO, validates the session, attaches the user’s EPM groups to the request
- Rate limits per user, blocks anonymous access, scrubs PII patterns before anything reaches the orchestrator
- L2 grounding reads dimension metadata from EPM on a schedule — not a hardcoded schema
- Only the schema + user query go to the model; financial values never leave the data layer
- L4 rejects anything outside the approved member lists and falls back to the deterministic parser
- Data stays inside the OCI boundary
- Natural fit when EPM is already in OCI
- Use the hyperscaler the org already governs
- Enterprise DPA, no training on prompts
- For regulated or sovereign data
- Highest control, highest run cost
- Oracle PCMCS — rule-set execution results for each method variant; the reciprocal solve runs in the PCMCS engine (iterative or matrix), not in browser algebra
- Least-privilege service account (read-only role, one app, one pod) with the token in a vault and rotated
- Results filtered to the requesting user’s EPM security before rendering
- Every query logged: user, timestamp, raw query, parsed intent JSON, model + prompt version, POV returned, latency, cost
- Exported to the SIEM; retained per the SOX evidence schedule
- Dashboards for fallback rate, eval pass rate, guardrail hits, p95 latency, spend
- The parsed JSON is shown to the user as the explanation (“AI: entity · year · scenario — 93% confident”) — the same line the demo prints today
- Every number on screen traces to an EPM cell intersection an auditor can reproduce
10.2 · Choosing the LLM platform
The demo’s DeepSeek call is a placeholder for a single adapter, callLLM(system, user), behind Layer 3. Swapping the provider changes one function and zero business logic. Pick the platform the organisation already governs — the security and procurement review is the long pole, not the integration.
| Option | Choose when | Data posture |
|---|---|---|
| Oracle OCI Generative AI (Cohere Command, Llama) | EPM Cloud already lives in OCI; you want one cloud boundary and one contract | Prompts stay in the OCI tenancy; no training on customer data; dedicated AI clusters available for isolation |
| Azure OpenAI Service | Microsoft-first finance estate (Entra ID, Purview, Sentinel already in place) | Private endpoint, regional deployment, zero-retention by default under the enterprise agreement |
| AWS Bedrock (Claude, Titan) / Google Vertex AI (Gemini) | The org’s landing zone is AWS or GCP; VPC endpoints and IAM already audited | VPC/PSC private access, no data used for training, CloudTrail/Cloud Audit Logs integration |
| Direct enterprise API (Anthropic, OpenAI) | Fastest model access; acceptable when a zero-data-retention agreement and DPA are signed | ZDR endpoint, SSO-managed keys, SOC 2 report on file |
| Self-hosted open weights (Llama, Mistral, Qwen via vLLM) | Sovereign or air-gapped requirements; regulated data classification forbids any external inference | Full control; you own patching, eval, and capacity — budget for an MLOps owner |
Put a model gateway in front of whichever you choose (Azure API Management, OCI API Gateway, Kong AI Gateway, LiteLLM, or Portkey): it owns key custody, per-team spend caps, routing and fallback between models, prompt/response logging, and lets you retire a deprecated model without touching the application.
10.3 · Security controls
| Control | Implementation |
|---|---|
| Identity & access | Covered in full in section 09 — corporate SSO, group-to-role mapping, and the decision about who enforces data-level security. Listed here because it is a production gate, not because it is optional. |
| Service account | One read-only EPM service account per application per pod, least-privilege role, no interactive login, credential in a vault (OCI Vault, Azure Key Vault, HashiCorp Vault), rotated on a schedule and on staff change. |
| Secrets & config | No secrets in code or build artifacts; environment-specific config injected at deploy; .dev.vars-style files never leave a developer machine. |
| Network | Private endpoints to the LLM provider and to EPM where the platform supports them; egress allow-list so the orchestrator can reach exactly two hosts; TLS 1.2+ everywhere. |
| Prompt-injection & input guardrails | Layer 1 (already in the demo) blocks instruction-override patterns, enforces length and scope; extend with a classifier on the gateway and log every rejection. |
| Output guardrails | Layer 4 (already in the demo) validates every returned member against the approved lists and strips unexpected keys; production adds a policy check that the resolved POV is inside the user’s security scope before the data call. |
| Data minimisation | Prompts contain metadata and the user’s query only. No cell values, no employee names, no free-text comments from EPM. Logged prompts are classified and retained accordingly. |
| Encryption | In transit (TLS) and at rest (provider-managed KMS); audit logs on immutable storage with customer-managed keys where policy requires. |
10.4 · SOX, audit, and model-risk controls
A read-only NLQ layer does not change a financial-reporting control, but it is an interface to a SOX-relevant system and lands squarely in ITGC scope. Treat prompts, schemas, and eval sets as code — that single decision satisfies most of what an auditor will ask for.
| Requirement | How it is satisfied |
|---|---|
| Complete, immutable audit trail | Append-only log of user, timestamp, raw query, parsed JSON, model and prompt version hash, POV returned, and row count — WORM storage, retained for the evidence period (typically 7 years), exported to the SIEM. |
| Change management | Prompt templates, few-shot examples, approved-member schema, and code are version-controlled; every change follows ticket → peer review → test evidence → CAB approval → deploy. A prompt edit is a code change. |
| Segregation of duties | Developers cannot deploy to production; the service-account owner is not a developer; production secrets are held by platform operations. |
| Access recertification | Quarterly review of who can use the tool and of the service account’s EPM roles, evidenced and signed. |
| Testing evidence | A golden-query regression suite (the few-shot examples plus a larger labelled set) runs in CI before every release; pass rate and diffs are archived as release evidence. |
| Model risk management | An inventory entry (intended use, limitations, owner, validation date) in the model-risk register — the SR 11-7 pattern for financial services; periodic re-validation when the model or prompt changes. |
| Reproducibility & lineage | Every displayed number traces to an EPM POV and a consolidation/calculation timestamp; an auditor can re-query the same intersection in EPM and match it. |
| Explainability | The parsed intent JSON is the explanation and is shown to the user on every response — no hidden reasoning between the query and the data call. |
10.5 · Dev → Test → Prod promotion
| Environment | EPM target | Data | Gate to leave |
|---|---|---|---|
| Dev | EPM Test pod (developer slice) | Synthetic or masked | Unit tests on the data layer; lint; eval suite ≥ threshold against the Test LLM deployment |
| Test / UAT | EPM Test pod (full refresh) | Masked copy of production | Business UAT sign-off on the golden queries; security scan; performance run (p95 latency, fallback rate) |
| Prod | EPM Production pod | Live | Change ticket approved; deploy in window; smoke test; hypercare with rollback ready |
- Promoted artifacts: application build, prompt templates (versioned), approved-member schema snapshot, eval set, infrastructure config (IaC) — all from the same Git tag.
- Pipeline: branch → PR review → CI (tests + evals) → deploy to Test → UAT sign-off → CAB → deploy to Prod → smoke test. Hosting can stay on Cloudflare Pages/Workers or move to OCI Functions + API Gateway or the org’s standard platform — the code does not care.
- Configuration: per-environment secrets and endpoints injected at deploy; the same build runs in every environment.
- Metadata sync: a scheduled job refreshes dimension metadata into Layer 2 grounding with change detection, so a new entity or account appears in the approved lists without a code release.
- Rollback: previous build and previous prompt version retained; rollback is a redeploy, and because prompts are versioned it also reverts a prompt regression.
10.6 · Operating it
- SLOs: p95 latency, availability of the read path (the deterministic fallback keeps it alive when the LLM is down — already built), fallback rate as a quality signal, eval pass rate per release.
- Cost governance: per-user and per-team token budgets at the gateway; alert on anomalies; the unit-cost model earlier in this kit is the baseline.
- Model lifecycle: providers retire models on a schedule — re-run the eval suite on the successor before switching, and record the switch as a change.
- Incident runbook: LLM outage → fallback parser; EPM API outage → cached metadata with a stale banner; guardrail spike → review logs for injection attempts.
10.7 · What changes for Reciprocal Costing
| Concern | Production answer |
|---|---|
| System of record | Oracle PCMCS — rule-set execution results for each method variant; the reciprocal solve runs in the PCMCS engine (iterative or matrix), not in browser algebra |
| Read/write posture | Read-only. Method selection is a costing-policy decision governed by change control, not a per-query toggle. |
| Use-case-specific control | Keep the side-by-side comparison for policy review only; the published cost-center numbers must come from the single approved method. |
10.8 · Production readiness checklist
- ✓LLM platform selected from the governed list, DPA / zero-retention terms on file, gateway in front of it
- ✓SSO integrated; authorisation derived from EPM security groups; cookie gate removed
- ✓Read-only EPM service account per pod, credential in a vault, rotation scheduled
- ✓Prompts, schema, and eval set version-controlled and under change management
- ✓Append-only audit log wired to the SIEM with the agreed retention
- ✓Golden-query eval suite passing in CI; results archived as release evidence
- ✓Dev / Test / Prod pipeline with gates, IaC, and a rehearsed rollback
- ✓Model-risk register entry and owner named; first re-validation date set
- ✓Metadata refresh job scheduled with change detection
- ✓SLOs, cost caps, and the incident runbook agreed with platform operations