GPT-6.1 Sol – OpenAI's New Generation Mainstream Model

Executive Summary:
GPT-6.1 Sol is a new generation mainstream model launched by OpenAI in 2026, serving as an upgraded version of GPT-6 Sol. It achieves nearly the same level of intelligence as the latter at just one-fi...
1. What is GPT-6.1 Sol
GPT-6.1 Sol is a new generation mainstream model launched by OpenAI in 2026, serving as an upgraded version of GPT-6 Sol. It achieves nearly the same level of intelligence as the latter at just one-fifth the cost per Token. This model focuses on three high-consumption scenarios: agent programming, computer operation, and professional work. It matches Astra on the DeepSWE v1.1 real-world code engineering tasks and lags only 2.1 percentage points behind in the OSWorld 2.0 computer operation evaluation. Additionally, it reduces cached input costs to just $0.10 per million Tokens. OpenAI has provided an Ultrafast mode with up to 8x speed, and the model is now available to paying users through API and products such as ChatGPT Work and Codex.

Technical Positioning and Domain: GPT-6.1 Sol belongs to the mid-tier inference models in the GPT-6 series. Its core positioning is to achieve a balance between cost and capability, serving as the "mainstream model" within OpenAI's product lineup. It continues the inference model architecture of the GPT-6 series, primarily targeting agent-based AI workloads, covering task types such as code generation, computer operation, and multi-step workflow automation, thus filling the market gap between the flagship model Astra and the entry-level models.
Development Background: This model was developed by OpenAI as an upgraded iteration of GPT-6 Sol. By employing post-training optimization, OpenAI has "scaled down" the capabilities of GPT-6 Astra to a lower cost range, while maintaining the inference model architecture. The focus of the post-training was on enhancing performance in three key scenarios: agent programming, computer operation, and professional tasks, enabling the mid-tier model to handle workloads that were previously only feasible with the flagship model.
Core Value: GPT-6.1 Sol addresses the pain point of balancing cost and effectiveness in agent-based AI applications. While maintaining intelligence levels close to the flagship model, it reduces standard input costs to $2 per million Tokens and cached input costs to $0.10, making long-running agent tasks economically viable. Its cached input cost is only one-tenth of Astra's, which holds practical significance for enterprise-level workflows involving multi-person collaboration.
Technical Features: The model supports multi-level inference intensity (Reasoning Effort) adjustments, ranging from low to high. It offers the Ultrafast high-speed mode, which provides up to 8 times the Token generation speed of the standard mode. Through the context caching reuse mechanism, agents can significantly reduce calling costs when reusing the same context across multiple requests. This is a key technical approach for cost optimization in this generation of models.
2. Key Features
Agent Programming Capabilities: On the DeepSWE v1.1 real-world codebase engineering tasks, the model matches GPT-6 Astra and outperforms GPT-6 Sol by 6.4 percentage points. The model can understand complex code repository structures, perform tasks such as bug fixing, feature implementation, and test completion, and output merge-ready Pull Requests, making it suitable as a resident coding agent.
Computer Operation Capabilities: In the OSWorld 2.0 computer usage evaluation, the model scores at 98.9% of Astra (2.1 percentage points behind), improving by 7 percentage points over GPT-6 Sol, while the single-task cost is approximately one-seventh of Astra's. The model can simulate human interaction with graphical interfaces to complete tasks such as file management, application interaction, and web browsing, serving as a core capability for desktop automation scenarios.
Professional Document Processing: In the GDP.pdf evaluation, the model scores 5.5 points higher than Claude Opus 5.5, with a single-task cost less than half of that of the latter. The model can parse tables, charts, and formatting information from complex PDF documents, extract structured data, and perform in-depth analysis, making it applicable to professional document scenarios such as contract review, financial report analysis, and academic paper interpretation.
Multi-step Business Workflow: On the AutomationBench evaluation, the model achieves a score 2.2 percentage points higher than Claude Opus 5.5 with medium inference intensity, at about one-third of the cost. This feature supports cross-system, multi-step business process automation, including data synchronization, form processing, and approval triggering, enabling it to handle enterprise-level workflow automation tasks.
Multi-level Inference Intensity Adjustment: The model supports adjustable inference effort levels, ranging from low to high. The low setting offers fast response times and low cost, suitable for simple tasks; the highest setting allocates more inference computational budget, ideal for complex code engineering and challenging professional tasks. The official reports scores on evaluations such as DeepSWE, OSWorld, and Terminal-Bench using the highest setting.
Ultrafast Generation Mode: Token generation speed can reach up to 8 times that of the standard version, enabled via API or a Pro subscription plan costing $500/month. This mode offers 8 times the speed for about 6 times the price compared to interactive programming scenarios like Codex, making it ideal for use cases sensitive to response latency.
Context Caching and Reuse Mechanism: The caching input price is set at $0.10 per million Tokens, which is 95% lower than standard input and 50% lower than GPT-6 Sol. When the agent reuses the same context across multiple requests, the server can reuse cached KV/prefix entries and only charge for the incremental input, significantly reducing the total cost of multi-turn interactions.
3. How to Use
Environment Requirements: GPT-6.1 Sol is available via the OpenAI API and the ChatGPT Work, Codex products, and does not require a local deployment environment. Users need to register for an OpenAI account and activate a paid plan — API users must bind a payment method, and ChatGPT users must subscribe to the corresponding paid tier of ChatGPT Work or Codex. It is recommended to activate the Pro tier with a monthly budget of $500 to access Ultrafast high-speed generation capabilities.
Using the ChatGPT Interface: Log in to ChatGPT (chatgpt.com), and in the ChatGPT Work or Codex interface, select the "GPT-6.1 Sol" option from the model dropdown menu. This model is not currently available in the standard Chat interface; users must switch to the Work or Codex workspace to use it. Once selected, you can begin conversations, write code, or execute work tasks.
API Calling Method: In your code, set the model ID to
gpt-6.1-soland initiate requests using the standard Chat Completions interface. Developers can use the official OpenAI SDKs (such as Python, Node.js) or directly call the REST API. The request body structure is consistent with other GPT-6 series models, resulting in low migration costs. Here is an example code snippet:
from openai import OpenAI
client = OpenAI()
response = client.chat.completions.create(
model="gpt-6.1-sol",
messages=[{"role": "user", "content": "Fix the memory leak issue in this Python script"}]
)
Setting Reasoning Effort: Choose the reasoning effort level based on the complexity of the task. For simple tasks (such as text classification, information extraction), use the low setting to save costs. For complex code engineering or multi-step reasoning tasks, use the high setting to achieve better results. The reasoning effort is controlled via the API parameter
reasoning_effort, with available values of low, medium, and high.Enabling Context Caching: Reuse system prompts and task context across multiple requests, so that the input portion of subsequent requests is billed at the cached rate of $0.10 per million Tokens, rather than the standard input rate of $2.00. It is recommended to place fixed instructions (such as system prompts, project background) at the beginning of the message sequence to allow the server to match the cache prefix.
Activating Ultrafast High-Speed Mode: When extremely fast generation is required, enable this mode by setting the
ultrafastparameter in the API request or by activating it in your ChatGPT Pro subscription. This mode provides a Token generation speed up to 8 times faster than the standard speed. Note that this mode uses a higher billing multiplier, so it is recommended to use it only in latency-sensitive scenarios (such as interactive coding, real-time translation).
4. Pros and Cons Analysis
| Pros |
|---|
| High cost-effectiveness: Standard input pricing is $2 per million Tokens, just one-fifth of the flagship Astra; cached input drops to $0.1, one-tenth of Astra's, making long-term operation costs for agent tasks manageable. |
| Strong agent programming capability: In the DeepSWE v1.1 real codebase evaluation, it matches GPT-6 Astra, improving by 6.4 percentage points over GPT-6 Sol, demonstrating the ability to handle real engineering tasks and complete the full workflow of bug fixing and PR submission. |
| Outperforms competitors in professional tasks: In the GDP.pdf evaluation, it surpasses Claude Opus 5.5 with a cost less than half of that model's; on AutomationBench, it leads Opus 5.5 by 2.2 percentage points at medium inference strength, showing excellent performance in professional scenarios. |
| Improved factual accuracy: Fact error rate decreases from 11.4% to 7.7% (a reduction of about 32%) at low inference strength, and further drops to 4.1% at high inference strength, approaching Astra's 4.0%, significantly enhancing reliability. |
5. Comparative Analysis with Similar Tools
| Comparison Dimension | GPT-6.1 Sol | GPT-6 Astra |
|---|---|---|
| Model Positioning | Mainstream model with the best balance of capability and cost | Flagship model with the strongest overall capability |
| Standard Input Cost (/ million Tokens) | $2.00 | $10.00 (5 times that of Sol) |
| Cached Input Cost (/ million Tokens) | $0.10 | $1.00 (10 times that of Sol) |
| DeepSWE v1.1 Programming Evaluation | Matches Astra | Baseline (highest) |
| OSWorld 2.0 Computer Operation | 2.1 percentage points behind Astra, single-task cost is about 1/7 | Baseline (highest) |
| GDP.pdf Professional Document Processing | Exceeds Opus 5.5, single-task cost is less than half | Industry-leading, cost is about 5 times that of Sol |
| Fact Error Rate (High-end) | 4.1% | 4.0% (slightly better) |
| Ultrafast High-speed Mode | Available, up to 8x speed | Available, fastest frontier model globally |
| Inference Intensity Adjustment | Supports low/medium/high multiple levels | Supports multiple levels |
Selection Recommendations: For agent programming, computer operations, and cost-sensitive long-running tasks, GPT-6.1 Sol is currently the most cost-effective option overall, especially suitable for SaaS products and enterprise automation workflows that require frequent API calls. Its cached input cost is only one-tenth of Astra's, giving it a clear advantage in agent workloads that involve multi-round context reuse.
For users seeking the ultimate capabilities in research reasoning and complex decision-making scenarios, GPT-6 Astra remains the best choice. It maintains the highest scores in scientific benchmarks such as Terminal-Bench Science, and also has a lower rate of search tool failure disclosures. If the user primarily handles professional documents and already has an existing dependency on the Claude ecosystem, Opus 5.5 is outperformed by Sol in GDP.pdf evaluations, but still holds its own advantages in other document understanding dimensions. The choice should be made based on specific business scenarios.
6. Editor's Summary
GPT-6.1 Sol represents a mature implementation of OpenAI's model product tiering strategy. From a technological innovation perspective, this model delivers flagship capabilities to the mid-tier price range through post-training optimization, rather than simply reducing model size. This approach effectively preserves the depth of the reasoning model while achieving systematic cost optimization through two mechanisms: cache reuse and inference intensity adjustment. The move to reduce input cache pricing to $0.1 directly alters the business model feasibility of agent-based AI applications.
In terms of practical value, GPT-6.1 Sol demonstrates real competitiveness in two core scenarios: intelligent agent programming and computer operation. The DeepSWE v1.1 evaluation results matching Astra indicate that development teams can confidently delegate common bug fixes and PR generation tasks to this model. Meanwhile, the OSWorld 2.0 score, which lags the flagship model by just 2.1 points, shows that desktop automation scenarios are already viable. Its cost advantages in professional documentation and business workflows make it an ideal choice for enterprise process automation.
The target audience for this model is clear: AI application developers, enterprise automation engineers, research data processors, and individual developers who frequently use coding assistants. For teams that need to control API budgets, GPT-6.1 Sol offers precise cost control capabilities through its low cache costs and adjustable inference levels. For interactive use cases that prioritize speed, the Ultrafast mode compensates for the performance limitations typically associated with mid-tier models.
Looking ahead, GPT-6.1 Sol validates the feasibility of the "flagship capability下沉" product strategy. In the future, OpenAI may introduce more Astra-level capabilities into the mid-tier price range, further widening the gap with open-source models. As context caching technology matures and inference intensity adjustment mechanisms improve, cost-efficient model tiering will become a standard product strategy for large model providers.
7. Application Scenarios
App Full Lifecycle Maintenance: After developers publish an app, a resident Agent continuously monitors user feedback channels. It automatically compiles frequently asked questions, identifies code defects, and fixes bugs. The Agent then submits a PR containing complete code changes and test results, while also generating a video explaining the changes for manual review. GPT-6.1 Sol's performance on the DeepSWE benchmark ensures its engineering capabilities in real-world codebases, and its low caching costs make long-running monitoring Agents economically viable.
Product Development Support: During the launch of a new product, the Agent continuously understands the product positioning and team requirements. After changes in requirements, it automatically updates release materials, adjusts product documentation, revises API descriptions, and marks sections requiring manual confirmation. The model's strengths on AutomationBench enable it to handle cross-document, multi-step information synchronization and material updates, reducing the administrative workload for product teams.
Scientific Research Data Analysis: After new experimental data is added to the database, the Agent automatically reruns the analysis process, investigates the causes of anomalies, updates paper charts, and synchronizes revisions to related textual explanations. GPT-6.1 Sol achieves a single-task cost of only $5.47 on the Terminal-Bench Science, which is 75% lower than the flagship model, offering a cost-effective automated data analysis solution for research teams, especially suitable for lab environments requiring frequent data processing iterations.
Enterprise Financial Automation: The Agent monitors accounts receivable status and automatically prepares invoice documents when invoices are missed, then sends them after manual confirmation. OpenAI has disclosed that early test users of Dot have successfully completed this task. This scenario validates the model's reliability in enterprise-level business processes. GPT-6.1 Sol's multi-step workflow capabilities and low operational costs make it well-suited as an automation assistant for financial workflows.
Multi-channel Information Monitoring: The Agent periodically checks specified Slack channels or email accounts, automatically aggregating new findings into a shared ChatGPT Space page and generating dynamic charts for team real-time viewing. The model's context caching mechanism ensures that long-running monitoring tasks only incur full input costs during initial loading, with subsequent incremental updates billed at $0.10 per million Tokens, significantly reducing ongoing monitoring operational costs.
8. FAQ
Q: What is the difference between GPT-6.1 Sol and GPT-6 Sol?
A: GPT-6.1 Sol is an upgraded version of GPT-6 Sol. It outperforms GPT-6 Sol by 6.4 percentage points in the DeepSWE v1.1 real-code benchmark and by 7 percentage points in the OSWorld 2.0 computer operation benchmark. Additionally, the input caching cost has been reduced by 50% from GPT-6 Sol, bringing it down to $0.10 per million Tokens, significantly improving overall cost-effectiveness.
Q: How should I choose between GPT-6.1 Sol and GPT-6 Astra?
A: The two models are positioned differently. GPT-6.1 Sol is the primary model, offering near-flagship intelligence levels at one-fifth the Token cost of Astra. It is ideal for cost-sensitive production environments and high-frequency API calls. GPT-6 Astra, on the other hand, is the most comprehensive flagship model, maintaining leadership in metrics such as scientific reasoning (highest score of 68.1% on Terminal-Bench Science) and search tool failure disclosure rate (1.5%). It is best suited for scenarios that demand extreme capabilities.
Q: What calling methods does GPT-6.1 Sol support?
A: The model can be accessed via the OpenAI API (model ID: gpt-6.1-sol) or selected within the ChatGPT Work and Codex product interfaces. It is not currently available in the standard Chat interface. To use its full capabilities, you must subscribe to the ChatGPT Work, Codex, or corresponding API paid plans.
Q: How can I use the Ultrafast mode when making API calls?
A: You can enable the Ultrafast mode by including the corresponding parameter in your API request, or by subscribing to the ChatGPT Pro plan at $500/month. Ultrafast mode provides up to 8 times the Token generation speed of the standard mode, but with a higher billing rate (approximately 6 times the price in Codex). It is recommended to enable this mode only in latency-sensitive scenarios such as interactive programming or real-time translation.
Q: How does context caching reduce costs?
A: When an agent reuses the same context across multiple requests, the OpenAI server can reuse the already cached KV/prefix and only charge for the incremental input added. The caching input cost is $0.10 per million Tokens, which is 95% lower than the standard input cost. It is recommended to place fixed system prompts and task background information at the beginning of the message sequence to improve cache hit rates.
Q: How should the inference strength adjustment be set for GPT-6.1 Sol?
A: The model supports multiple inference effort levels: low, medium, and high. Low effort provides fast response times and low cost, suitable for simple information extraction and text classification tasks. Medium effort is appropriate for general business workflows. High effort allocates more computational budget for inference and is suitable for complex code engineering and multi-step reasoning tasks. Official benchmark results are reported using the highest effort setting. In practice, it is recommended to dynamically adjust based on task complexity and budget.
9. Project Links
- OpenAI Official Website: https://openai.com/
- Official ChatGPT Entry: https://chatgpt.com/
- ChatGPT Pro Subscription Plan: https://chatgpt.com/plans/plus/
- OpenAI Technical Research and Publications: https://openai.com/research
- OpenAI API Documentation: https://platform.openai.com/docs
Related AI Model Articles

Claude Haiku 5.5 – Anthropic's Most Lightweight Model
Claude Haiku 5.5 is the latest small language model officially released by Anthropic on October 7, 2026. It is positioned as a high-value product with the tagline "cheapest, fastest, and most capable....
Ling-3.1-flash – A New Generation Large Model Launched by the Bailing Team at Ant Group
Ling-3.1-flash is a new generation large language model launched by the Bailing team at Ant Group. It employs a hybrid linear Attention architecture and MoE sparse activation technology, with a total ...

IQuest-Q1 Review: A 320B Sparse MoE Open-Source Agent Foundation Model Specializing in Code and Long-Horizon Agent Tasks
IQuest-Q1 is an open-source Agent foundation model developed by IQuestLab, with a core focus on code generation and Agent task execution. The model employs a sparse MoE architecture, featuring a total...

OpenViking – ByteDance's Open-Source AI Agent Context Database
OpenViking is an open-source AI Agent context database developed by ByteDance's Volcano Engine. Its core idea is to unify the Agent's memory, knowledge RAG, and skills within a custom `viking://` virt...
© All Rights Reserved. Some content on this site is partially generated by AI with human review.
