Last updated: August 26, 2026
Document-first. No account, no purchase, no workload testing. Every figure is read from a named vendor document with an access date. This briefing establishes published routing, pricing and retention behavior. It does not establish compliance with any regulation, and it cannot resolve anything that requires an authenticated API response or an invoice.
Amazon Bedrock exposes three routing options for GPT-5.6, but only two of them are cross-Region inference profiles. In-Region calls the raw model ID on the bedrock-mantle endpoint and keeps processing in one Region. Geographic and Global are system-defined profiles on bedrock-runtime. Geographic keeps processing inside a named geography. Global routes to any supported commercial AWS Region.
Verdict in one sentence: The 10 percent higher In-Region and Geographic rate buys a documented inference-processing boundary, and it does not by itself establish prompt-cache locality, which retention path a request enters, or whether your account can turn retention off.
| Fact | Value | Scope | Status |
|---|---|---|---|
| Regional rate | In-Region and Geo are 1.10x Global on all 24 published line items | Sol, Terra, Luna | Verified |
| Global rate vs OpenAI list | Identical on all 24 items | Standard tier | Verified |
| Geographic profile exists | 4 US Regions for all three models, plus Mumbai and Hyderabad for Terra and Luna | bedrock-runtime | Verified |
| Service tiers | Standard only. Priority, Flex and Reserved marked not supported | All three model cards | Official claim |
| Output-token burndown | 10x on bedrock-runtime | Sol, Terra, Luna | Official claim |
| Cache write vs read | 1.25x uncached input vs 0.10x, a 12.5x spread | GPT-5.6 family | Official claim |
| Prompt-cache physical location | No statement located | Search scope in Methodology | Not found in official sources |
| GPT-5.6 allowed retention modes | Not published in a static table | Requires authenticated call | Needs recheck |
| Default retention behavior | Two AWS pages describe it differently | Responses API | Contradicted |
| Sol price stability | Promotional at least through 21 November 2026 | OpenAI direct notice | Official claim |
All rows read from vendor documents on 26 August 2026. Status labels follow the FSR claim taxonomy.
- Workloads whose processing boundary is the United States, on any of the three models
- Workloads whose processing boundary is India, on Terra or Luna
- Workloads with no geography constraint that want the Global rate
- Buyers standardizing on AWS billing, IAM and logging
- A GPT-5.6 Geographic profile for the EU, Japan, Australia, Canada or the UK
- Where a GPT-5.6 prompt-cache entry is physically held
- Whether a given account can set retention mode to none for these models
- Whether any of this satisfies a specific legal obligation
In-Region is not a cross-Region profile
The three routing options are not three settings on one product. Two of them are inference profile IDs on bedrock-runtime, prefixed us., in. or global.. The third, In-Region, is the raw model ID called on bedrock-mantle, and the GPT-5.6 model cards mark In-Region as not supported on bedrock-runtime in every Region listed.
That distinction changes more than the destination. The two endpoints expose different feature sets for the same model. On bedrock-mantle, the cards list server-side tool calling, Projects and structured outputs as supported. On bedrock-runtime, server-side tool use, structured outputs, count tokens and application inference profiles are marked not supported, Projects is limited to the default project, and Guardrails is listed under the Converse API. AWS’s own launch blog describes server-side tool calling as available for all three models without naming an endpoint, so a reader who takes that sentence at face value will size the wrong architecture.
A procurement comparison that puts In-Region, Geographic and Global in one price column is comparing an endpoint change against a routing change.
Sources: AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026 · AWS, 20 August 2026
The ratio, and the denominator
Across the three GPT-5.6 model cards there are 24 published price points: three models, two context bands, four token line items each. On every one, the In-Region and Geographic rate is exactly 1.10 times the Global rate.
| Routing | Endpoint | Input | Cache write | Cache read | Output |
|---|---|---|---|---|---|
| In-Region | bedrock-mantle | $4.40 | $5.50 | $0.44 | $22.00 |
| Geo CRIS | bedrock-runtime | $4.40 | $5.50 | $0.44 | $22.00 |
| Global CRIS | bedrock-runtime | $4.00 | $5.00 | $0.40 | $20.00 |
The same 1.10x relationship holds for Terra, Luna and both context bands. Read from the AWS model cards on 26 August 2026.
Read in the other direction, Global is 9.09 percent below Regional. AWS uses that denominator and describes Global as offering approximately 10 percent savings, while stating that Geographic carries no surcharge for cross-Region routing. OpenAI uses the opposite denominator and describes regional processing endpoints as carrying a 10 percent uplift for models released on or after 5 March 2026. Both figures are correct. They are not the same number.
The Global column also matches OpenAI’s own Standard list price on all 24 items. That is arithmetic equality, and FSR does not treat it as evidence that one company copied the other’s policy. Public documents do not establish which structure came first.
Two facts stop this being a story about OpenAI specifically. xAI’s Grok 4.6 on the same Bedrock endpoint publishes the identical split, $2.20 and $6.60 for In-Region and Geographic against $2.00 and $6.00 for Global. And Bedrock does not apply a flat ratio across the catalog: Z.ai’s GLM 5 is $1.00 and $3.20 in the United States, $1.20 and $3.84 in several other Regions, and $1.55 and $4.96 in London.
Sources: AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026 · OpenAI, accessed 26 August 2026
Which geographies have a boundary
| Buyer requirement | Route established in the reviewed documents |
|---|---|
| One commercial US Region | In-Region on bedrock-mantle |
| Processing inside the United States | US Geographic profile, all three models |
| Processing inside India | India Geographic profile, Terra or Luna |
| Widest capacity, no geography constraint | Global profile |
| Processing inside the EU, Japan, Australia, Canada or the UK | No GPT-5.6 Geographic profile located as of 26 August 2026 |
| Documented prompt-cache locality | No statement located |
Availability finding tied to a date, not a legal conclusion. Read from the AWS regional availability table and model cards on 26 August 2026.
Workloads in Frankfurt, Tokyo, Sydney, London or Toronto are not blocked from GPT-5.6. They can call the Global profile. What they do not receive is a published geography boundary at any price.
Discovering that from the documentation is harder than it should be, because AWS enumerates its geographies four different ways: US, EU and APAC on the cross-Region inference overview; the prefixes us, eu and apac on the geographic profile page; US, EU, Japan and Australia in the prose of the regional availability page; and US, EU, APAC, JP and AU in the comparison table on that same page. None of the four names India, and India is the only non-US geography GPT-5.6 actually has.
Sources: AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026 · AWS, 18 August 2026
Where the post-launch expansions went
Sol has a US Geographic profile. What it does not have is a place in either constrained expansion AWS announced after the July launch. The India Geographic profiles announced on 18 August cover Terra and Luna. The GovCloud US-West and US-East launch announced on 24 August also names Terra and Luna. Both announcements repeat the same capability language AWS used at launch, including long-horizon genomics and biology analyses.
The specialist models sit at the other end of the same axis. Daybreak Red, the GPT-5.6 Cyber model, requires enrollment in OpenAI’s Trusted Access for Cyber program and runs in one Region, us-east-2, on bedrock-mantle, In-Region only. Its card publishes $13.75 input, $17.1875 cache write, $1.375 cache read and $82.50 output. Those are 1.10 times OpenAI’s direct rates for the same model. There is no Geographic or Global row on the card, so there is no cheaper routing option to decline.
Sources: AWS, 18 August 2026 · AWS, What’s New, 24 August 2026 · AWS, accessed 26 August 2026
Four data states, one published boundary
Routing controls where inference runs. It is not the only place GPT-5.6 content can come to rest, and the routing answer does not carry over to the others.
| Data state | Trigger | Published location | Duration |
|---|---|---|---|
| Inference processing | Every request | Set by the routing option | Request duration |
| Prompt cache | Repeated prefix over 1,024 tokens | No statement located | 30-minute TTL |
| Responses API state | store is true, which is the default under default mode | AWS retention path | At least 30 days |
| Safety retention | Classifier flag | Destination Region under CRIS | Up to 30 days |
Read from the AWS prompt caching, data retention and abuse detection pages on 26 August 2026.
Two of those rows need care. AWS’s abuse detection page states that Bedrock uses a zero data retention model and does not store inputs or outputs by default. AWS’s data retention page states that under the default mode the Responses API store parameter defaults to true, and adds that setting store=false does not guarantee zero retention because some models may still retain data for safety review. Those two descriptions of default behavior sit on the same documentation set and do not agree. A reviewer who reads only the first will build the wrong retention map.
The prompt cache is the other gap, and it is a priced one. Cache reads bill at a 90 percent discount to uncached input while cache writes bill at 1.25 times it, which makes a write 12.5 times a read and 25 percent more than not caching. AWS states that at times of high demand cross-Region routing optimizations may lead to increased cache writes, and GPT-5.6 caching runs in implicit mode unless you set it otherwise, placing an automatic breakpoint on the latest message. FSR searched the prompt caching, cross-Region inference, abuse detection and data retention pages and located no statement of where a cached prefix is physically held. That is a documented absence, not a claim that the cache leaves the boundary.
Zero data retention is a real setting. data_retention_mode: none is configurable at account and project level. What is not settled from public documents is whether GPT-5.6 permits it, because each model declares its own allowed modes and AWS does not publish those values in a static table. The resolution is an authenticated model call, and it belongs in the evidence file before procurement closes rather than after.
Sources: AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026
What the channel does not offer
All three GPT-5.6 model cards mark Priority, Flex and Reserved as not supported. Standard is the only tier.
The comparison is available on the same endpoint. Grok 4.6 publishes Priority at 1.75 times the Standard rate and Flex at 0.5 times. On a workload of one billion short-context Sol input tokens and one billion output tokens, the Global rate totals $24,000 and the Regional rate totals $26,400, so the routing decision is worth $2,400. A half-price tier at the multiplier Grok 4.6 publishes would be worth $12,000 on the same workload. GPT-5.6 does not have one on Bedrock, and OpenAI’s own platform lists Standard, Batch, Flex and Fast mode tabs for the same three models.
Quota behaves differently here too. On bedrock-runtime, GPT-5.6 output tokens burn quota at 10x, so one output token consumes ten from the tokens-per-minute allowance. Cache write tokens also count against it while cache read tokens do not. The Global cross-Region inference page carries an older burndown list that does not include GPT-5.6, so a team sizing a quota increase from that page alone will under-provision by an order of magnitude.
Sources: AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026 · AWS, accessed 26 August 2026
Contract jurisdiction is separate
AWS’s third-party model terms state that OpenAI Services on Amazon Bedrock are sold by OpenAI, excluding certain open-weight models and certain offerings available on AWS GovCloud, and that AWS is not a party to that agreement and has no liability or obligations under it. The agreement defines fees by reference to the Bedrock pricing page, which for the OpenAI frontier models presents a model selector rather than a rate table.
Its governing law clause splits customers in two. Those in the EEA, Switzerland and the UK are under the laws of Ireland with venue in Dublin. All other customers are under the laws of the State of California, with venue in San Francisco County. Processing location and contract jurisdiction answer different questions, and using the India Geographic profile does not change the second one.
One clause deserves a written answer rather than an interpretation. The agreement restricts integrating the service into products or services that provide biological or life sciences research and development, whether internally or for third parties, absent written approval from OpenAI or AWS, while explicitly permitting internal coding, software development and cybersecurity use. AWS’s launch material recommends Sol for drug discovery workflows. FSR offers no legal reading of that. The practical step is to ask both account teams, in writing and before procurement closes, whether the marketed workflow requires prior approval, who issues it, and what it covers.
Sources: AWS, version June 2026, accessed 26 August 2026 · AWS, 13 July 2026
The direct channel maps differently
Buying the same models from OpenAI produces a different footprint, and it is model-specific rather than purely regional.
OpenAI’s data controls guide separates regional storage from regional processing and lists them per region. Storage is offered in ten regions. Processing is marked yes for the United States, Europe, and the United Arab Emirates, and no for Australia, Canada, Japan, India, Singapore, South Korea and the United Kingdom. The UAE entry then narrows further: its processing row names specific model snapshots, and GPT-5.6 Luna appears among them. Every non-US region requires approval for abuse monitoring controls plus a Modified Retention amendment, and the guide notes that extended prompt caching in regions without regional processing may require content to be processed and temporarily stored outside the region.
Set against Bedrock, where the Geographic profile covers the United States and India, the two channels do not overlap the way most teams assume. An EU workload can obtain processing residency from OpenAI directly and not from Bedrock. An India workload can obtain it from Bedrock and not from OpenAI directly. A Japanese workload obtains it from neither in the documents reviewed here. Channel choice and residency choice are the same choice.
Sources: OpenAI, accessed 26 August 2026 · AWS, accessed 26 August 2026
FAQ
No EU Geographic profile for GPT-5.6 was located in the AWS documents reviewed on 26 August 2026. EU Regions can call the Global profile, which AWS defines as routing to any supported commercial Region worldwide. This is an availability finding tied to a date, not a compliance opinion.
No. Geographic and Global are inference profile IDs on bedrock-runtime. In-Region is the raw model ID on bedrock-mantle, and the GPT-5.6 cards mark In-Region as not supported on bedrock-runtime. The two endpoints expose different feature sets.
Geographic is exactly 1.10 times Global on all 24 published GPT-5.6 line items. Read in reverse, Global is 9.09 percent below Geographic. AWS describes the gap as approximately 10 percent savings on Global. OpenAI describes regional processing as a 10 percent uplift.
Not on the published record as of 26 August 2026. The India Geographic profiles cover Terra and Luna. Sol is reachable from Mumbai and Hyderabad through the Global profile, which routes worldwide.
AWS states that under the default retention mode the Responses API store parameter defaults to true, and that setting it false does not guarantee zero retention. A separate AWS page states that Bedrock does not store inputs or outputs by default. Treat the retention map as unresolved until you confirm your account’s effective mode.
Unknown from public documents. The mode exists and is set at account or project level, but each model declares its own allowed modes and AWS does not publish GPT-5.6’s values in a static table. Retrieve them from the model API and keep the response as evidence.
Standard only. All three model cards mark Priority, Flex and Reserved as not supported. Grok 4.6 on the same endpoint publishes Priority at 1.75x and Flex at 0.5x, which is the comparison worth putting to your account team.
Methodology
Document-first. FSR did not purchase, deploy, benchmark or observe GPT-5.6 on Amazon Bedrock, and no hands-on evidence appears above.
Sources are the AWS model cards for GPT-5.6 Sol, Terra, Luna, Daybreak Red and Grok 4.6, the Bedrock pricing page, the cross-Region inference, geographic and global profile, regional availability, prompt caching, abuse detection, data retention, token burndown and service tier pages of the Bedrock User Guide, the AWS third-party model terms dated June 2026, AWS What’s New announcements dated 17, 18 and 24 August 2026, the AWS launch blog dated 13 July 2026, and OpenAI’s developer pricing page and data controls guide. All were read on 26 August 2026 unless otherwise dated.
Where AWS pages disagree, both statements are preserved rather than reconciled. Three conflicts appear above: default retention behavior for the Responses API, the burndown list on the global cross-Region inference page, and the four enumerations of Bedrock geographies. Where a document was located but the relevant content could not be read, that is stated. Prompt-cache locality is recorded as not found in official sources after searching the four pages named in that section, which is not a claim that no boundary exists.
Several commercial AI research tools were used during preparation, including products from vendors named in this briefing. Any claim touching a tool vendor’s own products was re-verified from that vendor’s primary documentation through a separate route before use.
This is a Tier C briefing. FSR does not place affiliate links in Tier C briefings. This page carries none, and FSR receives nothing from an Amazon Bedrock, OpenAI or xAI subscription.
Verdict
The 1.10 ratio is exact and it is the smallest question here. It is not unique to OpenAI, it appears on a neighboring vendor’s model on the same endpoint, and on a large workload it is worth less than the discounted service tier GPT-5.6 does not have.
What the public record establishes is narrow. Bedrock documents a GPT-5.6 processing boundary in the United States for all three models and in India for two of them. It documents where classifier-flagged traffic is retained under cross-Region inference. It does not document where a prompt cache entry sits, it describes default retention two different ways on two of its own pages, and it does not publish the retention modes each GPT-5.6 model allows.
Teams whose boundary requirement resolves to the United States, or to India on Terra or Luna, have a documented route to evaluate. Everyone else is choosing a model whose geography profile has not been published for their region, and should retrieve the cache, retention mode and account-level answers in writing before that choice is signed rather than after.
FSR accepts paid buyer-side research engagements. The deliverable for this topic is a route and retention evidence matrix: destination Regions per profile, your account’s effective retention mode, the allowed modes each selected model returns, the guardrail enforcement path for your endpoint, service tier availability, and the written questions to put to your AWS and OpenAI account teams.
Contact FSRCommissioned work is disclosed on publication. Payment never changes a finding.
Tier B briefings include hands-on testing on a paid account. Tier C briefings are document-first, with no product testing.
Tier C briefing. Document-first, no hands-on testing, no product access. All figures read from named vendor documents on 26 August 2026. Pricing, regional availability and retention behavior for these models changed on 13 July, 30 July, 17 August, 18 August, 21 August and 24 August 2026, and should be re-verified before any procurement decision. GPT-5.6 Sol’s rate is described by OpenAI as promotional through at least 21 November 2026. FSR does not place affiliate links in Tier C briefings. This page carries none. FSR is not a law firm and gives no legal advice.