Back to Model List

IQuest-Q1 Review: A 320B Sparse MoE Open-Source Agent Foundation Model Specializing in Code and Long-Horizon Agent Tasks

AI Tech Editorial
RSS Feed
IQuest-Q1 Review: A 320B Sparse MoE Open-Source Agent Foundation Model Specializing in Code and Long-Horizon Agent Tasks official screenshot
(Image source: official screenshot)

Executive Summary:

IQuest-Q1 is an open-source Agent foundation model developed by IQuestLab, with a core focus on code generation and Agent task execution. The model employs a sparse MoE architecture, featuring a total...

1. What is IQuest-Q1

IQuest-Q1 is an open-source Agent foundation model developed by IQuestLab, with a core focus on code generation and Agent task execution. The model employs a sparse MoE architecture, featuring a total of 320B parameters, but only activates approximately 15B parameters per token. This design significantly reduces inference costs while maintaining flagship-level task capabilities. IQuest-Q1 has been trained using multi-harness reinforcement learning and multi-teacher online distillation (MOPD) in synthetic environments, allowing it to directly integrate with mainstream Agent tools such as Claude Code or Codex. It consistently ranks among the top open-source models on professional benchmarks like NL2Repo and CyberGym, offering a well-rounded performance with particular strengths in frontend generation, office automation, and AI research and development diagnostics.

iquest-q1-agent official website screenshot
(Image source: official screenshot)

Technical positioning and domain: IQuest-Q1 belongs to the intersection of natural language processing and Agent foundation models. Its unique positioning is as a model specifically designed for CLI Agent operations, repository-level code generation, and medium- to long-horizon multi-step task execution, rather than a generalized model focused on broad conversational abilities. It is intended for complex workflows in real-world production environments that require autonomous tool calling, code repository reading and writing, and recovery from errors.

Development background: This model was developed by IQuestLab, and during the research and development process, the team implemented a "human-AI collaborative self-evolving development process" — the model itself participates in capability diagnosis, training plan design, and partial experimentation, while human researchers are responsible for direction adjustment and high-cost experimental decisions. This development approach itself serves as a practical validation of the model's capabilities.

Core value: IQuest-Q1 addresses the contradiction in the open-source community between "strong code generation capabilities" and "low inference costs." By leveraging a sparse activation mechanism, it achieves performance at the level of a 320B parameter model with computational costs close to those of a 15B parameter model. Additionally, through MOPD multi-teacher online distillation, it consolidates multiple specialized reinforcement learning models into a single model, avoiding the resource waste associated with deploying multiple dedicated models.

Technical features: The model uses a decoder-only Transformer backbone combined with a sparse mixture-of-experts (MoE) structure, with training progressing through three stages: pre-training, mid-training, and post-training. Its differentiating advantage lies in multi-harness reinforcement learning — enabling a single strategy to be trained across multiple execution environments while preserving each environment's unique toolset and context management methods. This allows the model to learn not only problem-solving capabilities but also behavioral patterns for working within real Agent systems.

2. Key Features

  • Repository-level Code Generation: Users only need to input a single-sentence requirement, and the model can generate a complete, runnable multi-file project. On the NL2Repo leaderboard, IQuest-Q1 scores 63.0, ranking first among open-source models, surpassing DeepSeek-V4-Pro and Hun Yuan Hy4. This demonstrates its full-chain code generation capabilities from requirement understanding to engineering implementation.

  • CLI Agent Operations: The model can read and write to code repositories, invoke command-line tools, check execution results, and autonomously recover from errors. This capability allows it to seamlessly integrate with Claude Code and Codex, serving as a foundational reasoning engine to drive complete agent workflows, not just停留在对话式代码补全层面 (staying at the level of conversational code completion).

  • Frontend Interactive Application Generation: A single generation can produce a runnable web application that includes 3D scenes, interactive logic, and visual presentation. Typical examples include FPS games, racing games, and voxel-based underwater cities. This feature covers the complete frontend development pipeline, from scene setup, interactive logic, to visual rendering.

  • Full Office Workflow Automation: By cross-referencing chat records and documents using the lark-cli from Feishu, the model can complete a full cycle of event classification, report writing, task assignment, and progress tracking. This capability extends the Agent from the "code world" to the "office world," achieving cross-application information integration and workflow automation.

  • AI R&D Self-Diagnosis: The model can trace anomalies in the reinforcement learning training process, identify root causes, fix the code, and validate the effectiveness of the fix using new metric curves. This feature highlights the model's meta-level application value — it not only assists humans in developing AI, but also participates in the R&D iteration of AI itself.

  • Multi-step Complex Task Execution: Designed for medium- to long-range tasks, the model can autonomously advance complex workflows that last for several hours, covering scenarios such as network security and terminal operations. It consistently ranks among the top performers on eight benchmarks including DeepSWE, Terminal-Bench, and JobBench, with no significant capability shortcomings.

3. How to Use

  1. Obtain model weights: Download the model weights and inference code from Hugging Face or the GitHub repository. The Hugging Face model repository is located at https://huggingface.co/IQuestLab/IQuest-Q1, and the GitHub repository is at https://github.com/IQuestLab/IQuest-Q1. The ModelScope channel will be available soon.

  2. Deploy inference service: Launch an OpenAI-compatible inference service using SGLang or vLLM. The hardware requirements include 8-card tensor parallelism (TP=8) and bfloat16 precision. Additionally, you must load the IQuest-specific tool-call and reasoning parsers to ensure correct parsing of tool calls and reasoning processes.

  3. Acceleration and parameter tuning (optional): If you require higher throughput, enable the recursive MTP speculative decoding feature. To reproduce the official evaluation results, it is recommended to set temperature=1.0, top-p=0.95, and top-k=20. These parameter combinations have been verified by the official team to achieve benchmark performance.

  4. Direct model invocation: Initiate requests directly via the Python API or an OpenAI-compatible interface. Use the IQuest chat template format for tool calls and reasoning to ensure the request format aligns with the data distribution used during model training.

  5. Integrate with Agent tools: Connect to Claude Code or Codex CLI through a gateway that supports tool calling. After configuring the gateway address and API Token, you can use IQuest-Q1 as the underlying model to drive agents for code generation, tool calling, and long-range task execution.

4. Pros and Cons Analysis

Pros
Exceptional Inference Efficiency: With a total of 320B parameters but only 15B activated per token, it achieves flagship-level task capabilities with near-small-model inference costs, significantly reducing computational expenses while maintaining performance.
Balanced and Comprehensive Performance: It consistently ranks at the top across eight benchmarks, including DeepSWE, NL2Repo, CyberGym, Terminal-Bench, and JobBench, with no obvious performance weaknesses, making it suitable for unified deployment across multiple scenarios.
Outstanding Code Generation Ability: It achieves a score of 63.0 on NL2Repo for repository-level generation, surpassing DeepSeek-V4-Pro and HunYuan Hy4, placing it among the first-tier open-source models, with strong engineering implementation capabilities.
Real Intelligent Agent Loop: It can not only generate code but also read logs, call tools, and recover from errors, delivering verifiable final results rather than just停留在文本生成层面 (stopping at the text generation level).
One-Click Integration with Major Agents: It provides an OpenAI-compatible interface and an official Docker image, allowing direct replacement in Claude Code and Codex with low migration costs and good ecosystem compatibility.

5. Comparative Analysis with Similar Tools

Comparison Dimension IQuest-Q1 (ZhiZhi Innovation) DeepSeek-V4.1-Flash (DeepSeek) DeepSeek-V4-Pro (DeepSeek)
Core Architecture Sparse MoE, decoder-only Transformer, total parameters 320B / activated 15B Sparse MoE, Flash is the efficient inference variant Sparse MoE, flagship configuration
Performance Metrics (NL2Repo) 63.0 (first-tier open-source) Not listed 61.5
Performance Metrics (DeepSWE v1.1) 64.3 74.2 Not disclosed
Performance Metrics (CyberGym) 84.5 88.1 Not disclosed
Feature Highlights Specialized in CLI agents, repository-level code generation, multi-harness reinforcement learning, MOPD distillation Lightweight inference version of a general-purpose flagship model, focusing on high throughput and low latency General-purpose flagship model, covering a wide range of NLP tasks
Deployment Method 8-card TP=8, SGLang/vLLM, Docker image Lightweight deployment, low resource requirements High-performance cluster deployment
Open Source License Open source (weights available for download) Open source Open source
Community Ecosystem New project, community in early stages Mature community, well-established ecosystem Mature community, well-established ecosystem

Selection Recommendations: If your primary use case involves CLI agent operations, repository-level code generation, or requires integration with Claude Code / Codex-based agent workflows, IQuest-Q1's leading performance on NL2Repo and its specialized focus make it the preferred choice among open-source models. Its architecture design of 320B total parameters / 15B activated provides a good balance between inference cost and task capability.

If you prioritize general-purpose inference capabilities and a mature community ecosystem, DeepSeek-V4.1-Flash achieves higher scores on DeepSWE and CyberGym. The Flash version is optimized for efficient inference, making it suitable for large-scale production environments with high throughput requirements. For teams deeply integrated with the Tencent Cloud ecosystem, Hy4's cloud-native integration advantages are worth considering, although its performance on specialized code generation benchmarks falls short of IQuest-Q1.

6. Editor's Summary

IQuest-Q1 represents a noteworthy engineering practice in the open-source Agent foundation model domain. The core of its technological innovation lies in two aspects: first, a multi-harness reinforcement learning strategy that trains the model in an executable environment composed of real APIs, MCP servers, and repositories, with failure attribution performed before optimization to ensure that only the model's own errors serve as learning signals—this design directly enhances the behavioral reliability of the model in real agent systems. Second, the MOPD multi-teacher online distillation approach folds four specialized reinforcement learning models into a single model, addressing the deployment challenge of "specialist models being difficult to integrate."

In terms of practical value, IQuest-Q1's performance on NL2Repo that surpasses DeepSeek-V4-Pro is a milestone, proving that open-source models now have the capability to compete with commercial closed-source models on the high-difficulty task of repository-level code generation. Its "one-click integration with Claude Code and Codex" design reduces migration costs, while the "AI R&D self-diagnosis" feature highlights the model's potential for meta-level applications.

In terms of target users, this model is suitable for AI teams with an 8-GPU card cluster configuration, developers of agent applications, and enterprise technical departments requiring automated office workflows. For individual developers or teams with limited computational resources, the 15B activation parameter architecture implies that its inference cost is lower than that of dense models with comparable capabilities. However, the feasibility of hardware investment still needs to be evaluated.

Looking ahead, the "human-machine collaborative self-evolving R&D process" of IQuest-Q1 is worth noting—the model participates in diagnosing its own capabilities and designing training strategies. If this mode continues to iterate, it could potentially accelerate the evolution speed of model versions. However, as a newly open-sourced project, its community ecosystem, documentation completeness, and third-party toolchain still require time to mature. The listing progress on the ModelScope platform is also something to watch.

7. Application Scenarios

  • Intelligent Programming Development: Developers can input a single-sentence requirement, and IQuest-Q1 can generate a complete, runnable multi-file project. In environments like Claude Code or Codex, the model can autonomously read code, fix bugs, run tests, and deliver verifiable engineering results, making it suitable for rapid prototyping and codebase maintenance.

  • Interactive Web and Game Generation: Based on textual descriptions, the model can generate frontend applications such as 3D games, visualization tools, and design pages in one go. This scenario covers the complete workflow from scene setup, interactive logic, to visual presentation, ideal for game prototyping, data visualization, and creative showcases.

  • Office Process Automation: By integrating with the lark-cli of Feishu, the model connects chat, documents, spreadsheets, and meetings to complete the entire process of cross-verification of information, event classification, report writing, task assignment, and progress tracking. It is applicable to enterprise internal operations management, project collaboration, and event response scenarios.

  • AI R&D and Training Operations: The model can automatically diagnose training anomalies, identify root causes, and fix code, assisting AI teams in maintaining and optimizing their own R&D pipelines. It is suitable for scenarios such as reinforcement learning training monitoring, model hyperparameter tuning, and experimental validation, reducing the cost of manual troubleshooting.

  • Cybersecurity and Terminal Operations: For medium- to long-term tasks, the model can autonomously advance for several hours, covering scenarios such as cybersecurity detection and terminal command execution. Its performance of 84.5 on the CyberGym leaderboard validates its usability in security attack-defense simulations and system operations.

8. FAQ

Q: What are the hardware deployment requirements for IQuest-Q1?
A: The official recommendation is to use an 8-card tensor parallel (TP=8) configuration with bfloat16 precision, and the inference framework supports SGLang or vLLM. The model has a total of 320B parameters, but only 15B are activated per token. Therefore, the memory usage during inference mainly focuses on loading the model weights, and it is necessary to ensure that the memory of a single GPU is sufficient to hold the corresponding weight shards.

Q: How to choose between IQuest-Q1 and DeepSeek-V4-Pro?
A: In the NL2Repo repository-level code generation task, IQuest-Q1 outperforms DeepSeek-V4-Pro with a score of 63.0 compared to 61.5, showing a stronger advantage in code generation capabilities. However, in tasks such as DeepSWE and CyberGym, DeepSeek-V4.1-Flash achieves higher scores. If the core scenario involves agent-based code manipulation, choose IQuest-Q1; if the goal is general-purpose inference and higher throughput, consider the DeepSeek series.

Q: How to integrate IQuest-Q1 with Claude Code or Codex?
A: Through a gateway that supports tool calling. Users need to configure the gateway address and API Token. IQuest-Q1 provides an OpenAI-compatible interface, allowing direct replacement of existing model configurations. Additionally, the official provides a Docker image to simplify the deployment process.

Q: Does the model support Chinese?
A: The original information indicates that the model was trained in the middle stage to shift the data distribution toward code and STEM, and long-context agent trajectories were added. However, it does not explicitly mention any specialized optimization for Chinese capabilities. It is recommended that users in Chinese scenarios conduct targeted testing before actual deployment to confirm the model's performance.

Q: What is the training process for IQuest-Q1?
A: The training is divided into three stages: pre-training establishes general capabilities and world knowledge; mid-training shifts the data distribution toward code and STEM, incorporating long-context agent trajectories; and post-training focuses on specialized SFT (Supervised Fine-Tuning) and reinforcement learning convergence for software engineering, long-horizon tasks, and general inference. The post-training stage employs multi-harness reinforcement learning and MOPD (Multi-Teacher Online Distillation) with multiple teachers.

Q: How to reproduce the official evaluation results?
A: The official recommends setting temperature=1.0, top-p=0.95, and top-k=20, and loading the IQuest-specific tool-call and reasoning parsers. Enabling recursive MTP speculative decoding can improve throughput, but it may affect the precise reproduction of evaluation results.

9. Project Links

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.