OutYet reporting

GPT-5.6 Sol gains an Australian Bedrock route, with global routing caveats

AWS has documented a Bedrock path for GPT-5.6 Sol from Sydney and Melbourne, but the global profile can process requests outside those source Regions.

OutYet Editorial Desk

AWS published a September 2 guide describing access to GPT-5.6 Sol through Amazon Bedrock from the Asia Pacific (Sydney) and Asia Pacific (Melbourne) Regions. The guide names the global inference profile `global.openai.gpt-5.6-sol`, alongside Terra and Luna profiles. This is a regional Bedrock-access update, not evidence from OutYet's detector that establishes or changes the model's release state.

The regional entry point should not be confused with an in-country processing guarantee. AWS says an application sends its request to the Bedrock Runtime endpoint in Sydney or Melbourne, then Bedrock routes it to a supported commercial AWS Region. Its cross-Region inference documentation likewise says a global profile automatically selects a commercial Region for processing, so teams with data-residency requirements need to inspect the routing model rather than infer residency from the endpoint they call.

For teams already using OpenAI-shaped clients, AWS documents three invocation paths: the OpenAI Responses API, OpenAI Chat Completions API, and Bedrock's Converse API. The OpenAI-compatible routes use Bedrock Runtime's `/openai/v1` endpoint and can authenticate with SigV4 or a Bedrock model inference API key. That gives existing applications a migration path, but it also makes AWS identity, permissions, and endpoint configuration part of the integration.

The guide also gives Codex users a concrete deployment pattern: configure Codex with the Bedrock Runtime provider, point it at `global.openai.gpt-5.6-sol`, and use an AWS profile backed by an OIDC credential process. AWS says the helper exchanges an identity-provider token for temporary AWS credentials, after which requests are signed with SigV4. That pattern may suit organizations that want inference access governed through AWS federation instead of distributing a static inference credential.

Capacity and quota behavior remain operational constraints. AWS says GPT-5.6 on-demand quotas are measured in requests per minute and tokens per minute, with output tokens consuming quota at a higher burndown rate than input tokens. It also warns that inference-profile membership and model availability can change, and recommends checking support, requesting quota increases early, and testing representative prompts, output lengths, streaming, concurrency, and peak traffic before production rollout.

Related models

Sources