View a markdown version of this page

GPT-6.1 Sol - Amazon Bedrock

GPT-6.1 Sol

OpenAI logo. OpenAI — GPT-6.1 Sol

Model Details

OpenAI GPT-6.1 Sol brings advanced capabilities to workloads where both performance and cost matter. GPT-6.1 Sol helps agents investigate codebases and iterate on solutions. Agents can also use it to understand complex documents and complete business and computer use workflows across multiple steps.

  • Model launch date: September 29, 2026

  • Model lifecycle policy: OpenAI model deprecation notice periods. This model follows OpenAI first-party lifecycle terms, with at least 6 months of deprecation notice for generally available models, unless safety or compliance concerns require a faster timeline.

  • Model EOL date: Not announced.

  • End User License Agreements and Terms of Use: OpenAI models on Amazon Bedrock terms

  • Model lifecycle: Active

  • Context window: 1M tokens

  • Max output tokens: 131,072 tokens

  • Marketplace product ID: prod-qco655ut2vn54

Input Modalities Output Modalities
not-supported Audionot-supported Embedding
supported Imagenot-supported Image
not-supported Speechnot-supported Speech
supported Textsupported Text
not-supported Videonot-supported Video

Endpoints and APIs supported

The following tables show which endpoints and APIs GPT-6.1 Sol supports. For more information, see APIs supported by Amazon Bedrock and Endpoints supported by Amazon Bedrock.

Endpoint support

EndpointSupported
bedrock-runtimesupported
bedrock-mantlesupported

APIs supported on bedrock-runtime

MessagesResponsesChat CompletionsConverseInvoke
not-supported supported supported supported supported

APIs supported on bedrock-mantle

MessagesResponsesChat CompletionsConverseInvoke
not-supported supported supported not-supported not-supported
Note

On bedrock-mantle, both APIs use the /openai/v1 base path, not /v1. Use either API with this model:

  • For Responses, use /openai/v1/responses.

  • For Chat Completions, use /openai/v1/chat/completions.

Capabilities and Features

Bedrock Features

Features supported using bedrock-runtime

SupportedNot Supported

Features supported using bedrock-mantle

Prompt caching

GPT-6.1 Sol supports both Implicit Prompt Caching and Explicit Prompt Caching. API support differs by endpoint:

  • Responses and Chat Completions support both caching types on the bedrock-runtime and bedrock-mantle endpoints.

  • InvokeModel supports both caching types on the bedrock-runtime endpoint.

  • Converse supports implicit caching on the bedrock-runtime endpoint. The native cachePoint field isn't supported for explicit caching with this model.

For explicit caching, set prompt_cache_options.mode to explicit. Set prompt_cache_options.ttl to 30m, the only supported TTL and the default. In Responses requests, add "prompt_cache_breakpoint": {"mode": "explicit"} to an input_text content block. In Chat Completions and InvokeModel requests, add it to a text content part.

A cacheable prompt prefix must contain at least 1,024 tokens. You can include multiple explicit breakpoints, and each request can create up to four cache writes. Check the cache usage fields in the response to determine whether tokens were written to or read from the cache.

JSON Schema output

Use JSON Schema to set the format of the model response for streaming and non-streaming requests. Chat Completions and Responses support JSON Schema output on both the bedrock-runtime and bedrock-mantle endpoints. InvokeModel, InvokeModelWithResponseStream, Converse, and ConverseStream support JSON Schema output on the bedrock-runtime endpoint. Choose a model ID or inference profile ID from Programmatic access.

  • For Chat Completions, set response_format.type to json_schema. Put name, schema, and strict: true in response_format.json_schema.

  • For Responses, set text.format.type to json_schema. Put name, schema, and strict: true in text.format.

  • For InvokeModel and InvokeModelWithResponseStream, use the same response_format fields as Chat Completions.

  • For Converse and ConverseStream, set outputConfig.textFormat.type to json_schema. In outputConfig.textFormat.structure.jsonSchema, set name and schema. Encode the schema as a JSON string. Then set additionalModelRequestFields.text.format.strict to true.

Use an object schema. Put all fields in required. Set additionalProperties to false. Validate your schema before you send it. Check for refusals or incomplete responses before you parse the output.

Important

For JSON Schema output with Converse and ConverseStream, you must set additionalModelRequestFields.text.format.strict to true, in addition to specifying the schema in outputConfig. Include the following field in your request.

{ "additionalModelRequestFields": { "text": { "format": { "strict": true } } } }

Pricing

All prices are in USD per 1 million tokens. The following tables list Standard and Ultrafast prices. Global CRIS rates match OpenAI first-party pricing for the same service tier.

In-Region access on bedrock-mantle costs 10% more than Global CRIS. US geographic cross-Region inference (US CRIS) on bedrock-runtime has the same premium. This applies to both service tiers. The tables include this premium.

GPT-6.1 Sol supports both implicit and explicit prompt caching. Cache-write tokens are billed at 1.25× the uncached input-token rate, and cache-read tokens are billed at 0.05× the uncached input-token rate. For configuration and API support, see Prompt caching.

Long-context rates apply to the full request when input exceeds 272,000 tokens.

Priority and Flex tiers are not supported for this model.

Standard — Commercial Regions, short context (272K input tokens or fewer)

Inference optionInputInput — cache writeInput — cache readOutput
Regional (bedrock-mantle in US East (N. Virginia)) $2.20 $2.75 $0.11 $11.00
US CRIS (bedrock-runtime) $2.20 $2.75 $0.11 $11.00
Global CRIS (bedrock-runtime) $2.00 $2.50 $0.10 $10.00

Standard — Commercial Regions, long context (more than 272K input tokens)

Inference optionInputInput — cache writeInput — cache readOutput
Regional (bedrock-mantle in US East (N. Virginia)) $4.40 $5.50 $0.22 $16.50
US CRIS (bedrock-runtime) $4.40 $5.50 $0.22 $16.50
Global CRIS (bedrock-runtime) $4.00 $5.00 $0.20 $15.00

Ultrafast costs six times as much as Standard. Global CRIS rates match OpenAI first-party Ultrafast pricing. On bedrock-runtime, use Ultrafast through US CRIS or Global CRIS. On bedrock-mantle, use Ultrafast in us-east-1. Short-context prices apply to requests with up to 272,000 input tokens. Above that limit, long-context prices apply to the full request.

Ultrafast — Commercial Regions, short context (272K input tokens or fewer)

Inference option Input Input — 30m cache write Input — cache read Output
In-Region (bedrock-mantle in us-east-1)$13.20$16.50$0.66$66.00
US CRIS (bedrock-runtime)$13.20$16.50$0.66$66.00
Global CRIS (bedrock-runtime)$12.00$15.00$0.60$60.00

Ultrafast — Commercial Regions, long context (more than 272K input tokens)

Inference option Input Input — 30m cache write Input — cache read Output
In-Region (bedrock-mantle in us-east-1)$26.40$33.00$1.32$99.00
US CRIS (bedrock-runtime)$26.40$33.00$1.32$99.00
Global CRIS (bedrock-runtime)$24.00$30.00$1.20$90.00

Programmatic Access

To call this model from code, use the following model IDs and endpoint URLs. For more information, see APIs supported by Amazon Bedrock and Endpoints supported by Amazon Bedrock.

EndpointModel IDIn-Region endpoint URLGeo inference IDGlobal inference ID
bedrock-mantleopenai.gpt-6.1-solhttps://bedrock-mantle.us-east-1.api.aws/openai/v1Not supportedNot supported
bedrock-runtimeopenai.gpt-6.1-solNot supportedus.openai.gpt-6.1-solglobal.openai.gpt-6.1-sol

For in-Region access, use bedrock-mantle in us-east-1 (N. Virginia). On bedrock-runtime, the OpenAI-compatible base URL is https://bedrock-runtime.{region}.amazonaws.com/openai/v1. Use us.openai.gpt-6.1-sol for US geographic cross-Region inference or global.openai.gpt-6.1-sol for global cross-Region inference. Direct in-Region invocation is not supported on bedrock-runtime. Use a source Region enabled for the profile you choose; see Route model inference requests across AWS Regions with cross-Region inference.

Service Tiers

Amazon Bedrock offers several service tiers for different workloads. Standard gives you pay-per-token access with no commitment. To use it, set "service_tier": "default" or omit the field. For more information, see service tiers.

Ultrafast is a premium speed tier for GPT-6.1 Sol. Use it for workloads where response speed matters most. To request it with the Responses API, set "service_tier": "ultrafast".

StandardUltrafastPriorityFlexReserved
supported supported not-supported not-supported not-supported

Use one of these routes for Ultrafast:

  • bedrock-mantle in us-east-1, with model ID openai.gpt-6.1-sol.

  • bedrock-runtime, with us.openai.gpt-6.1-sol for US CRIS or global.openai.gpt-6.1-sol for Global CRIS.

Regional Availability

Regional availability at a glance

Standard and Ultrafast are available on bedrock-mantle in us-east-1 (N. Virginia). Both tiers support US geographic and global inference profiles on bedrock-runtime. For more information, see Regional availability by models.

Availability using the bedrock-mantle endpoint

RegionIn-RegionGeoGlobal
us-east-1 (N. Virginia)supportednot-supportednot-supported

Availability using the bedrock-runtime endpoint

ScopeIn-RegionGeoGlobal
US geographic and global inferencenot-supportedsupportedsupported

Geo inference details

The destination Regions available to a geographic inference profile depend on the source Region. To retrieve the current routing configuration, call GetInferenceProfile from the source Region.

Geo: US

Geo inference ID: us.openai.gpt-6.1-sol

Source Region Destination Regions
us-east-1 (N. Virginia)us-east-1 (N. Virginia), us-east-2 (Ohio), us-west-2 (Oregon)
us-east-2 (Ohio)us-east-1 (N. Virginia), us-east-2 (Ohio), us-west-2 (Oregon)
us-west-1 (N. California)us-east-1 (N. Virginia), us-east-2 (Ohio), us-west-1 (N. California), us-west-2 (Oregon)
us-west-2 (Oregon)us-east-1 (N. Virginia), us-east-2 (Ohio), us-west-2 (Oregon)
ca-central-1 (Canada)ca-central-1 (Canada), us-east-1 (N. Virginia), us-east-2 (Ohio), us-west-2 (Oregon)
ca-west-1 (Calgary)ca-west-1 (Calgary), us-east-1 (N. Virginia), us-east-2 (Ohio), us-west-2 (Oregon)

Quotas and Limits

Quotas vary by account and Region. See Quotas for Amazon Bedrock and the model's service quota settings.

The output-token burndown rate is 10: each output token consumes 10 tokens of quota.

Sample Code

Step 1 - AWS Account: If you already have an AWS account, skip this step. If you are new to AWS, sign up for an AWS account.

Step 2 - API key: Go to the Amazon Bedrock console and generate a long-term API key.

Step 3 - Get the SDK: You must have Python installed to use this guide. Then install the OpenAI SDK.

python3 -m pip install openai

Step 4 - Set environment variables

Set up your environment to use the API key for authentication.

bedrock-mantle
export OPENAI_API_KEY="<provide your Bedrock API key>" export OPENAI_BASE_URL="https://bedrock-mantle.us-east-1.api.aws/openai/v1"
bedrock-runtime
export OPENAI_API_KEY="<provide your Bedrock API key>" export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"
Note

On bedrock-runtime, set the model to us.openai.gpt-6.1-sol for US CRIS or global.openai.gpt-6.1-sol for Global CRIS. Direct in-Region invocation is not supported on this endpoint.

Step 5 - Run your first inference request

Save the file as bedrock-first-request.py.

bedrock-mantle

Use the settings from Step 4 - Set environment variables. Choose the bedrock-mantle tab. The Chat tab uses the Chat Completions API.

Responses API
from openai import OpenAI client = OpenAI() response = client.responses.create( model="openai.gpt-6.1-sol", input="Can you explain the features of Amazon Bedrock?", max_output_tokens=512, ) print(response.output_text)
Chat
from openai import OpenAI client = OpenAI() response = client.chat.completions.create( model="openai.gpt-6.1-sol", messages=[{"role": "user", "content": "Can you explain the features of Amazon Bedrock?"}] ) print(response.choices[0].message.content)

bedrock-runtime: OpenAI SDK

Use the settings from Step 4 - Set environment variables. Choose the bedrock-runtime tab. Send your request with the Responses API. The example uses the US CRIS profile. For Global CRIS, replace us.openai.gpt-6.1-sol with global.openai.gpt-6.1-sol.

Responses API
from openai import OpenAI client = OpenAI() response = client.responses.create( model="us.openai.gpt-6.1-sol", input="Can you explain the features of Amazon Bedrock?", max_output_tokens=512, ) print(response.output_text)

Ultrafast example

Set OPENAI_API_KEY to your Amazon Bedrock API key. This example uses the bedrock-mantle endpoint in us-east-1.

export OPENAI_API_KEY="<your Amazon Bedrock API key>"
from openai import OpenAI client = OpenAI( base_url="https://bedrock-mantle.us-east-1.api.aws/openai/v1", ) response = client.responses.create( model="openai.gpt-6.1-sol", service_tier="ultrafast", input="Explain how Amazon Bedrock cross-Region inference works.", ) print(response.output_text)

For bedrock-runtime, use the OpenAI-compatible base URL from Programmatic Access. Set the model to us.openai.gpt-6.1-sol for US CRIS or global.openai.gpt-6.1-sol for Global CRIS, and keep service_tier="ultrafast".