In-Depth Review of Startlux-Decision – StartLux Labs' Open-Source Decision Model Family
Executive Summary:
Startlux-Decision is an open-source decision model family introduced by StartLux Labs, covering five dense models ranging from 0.8B to 27B, as well as a 35B-A3B mixed expert model. It is specifically ...
1. What is Startlux-Decision
Startlux-Decision is an open-source decision model family introduced by StartLux Labs, covering five dense models ranging from 0.8B to 27B, as well as a 35B-A3B mixed expert model. It is specifically designed for typed decision-making tasks. This model family accepts state inputs in the form of text, JSON, or images, and directly returns the probability distribution of each option through a single forward inference for structured questions such as selection, yes/no, and scoring. This makes it naturally suitable for use as confidence levels. The model natively supports ultra-long context of up to 256K tokens, has millisecond-level inference latency, and is compatible with Jev's TypeSafe client format, providing an efficient and quantifiable technical foundation for automated decision-making and agent applications.
Technical Positioning and Domain: Belongs to the intersection of natural language processing and decision intelligence, focusing on shifting the capabilities of large language models from open-ended generation to structured, quantifiable decision outputs. Unlike general-purpose conversational models, Startlux-Decision is optimized at the architectural level for the "read state – output probability" paradigm, making it suitable for automated scenarios requiring determinism, explainability, and risk quantification.
Development Background: Developed by the StartLux Labs team, which has deep expertise in efficient inference, linear attention mechanisms, and MoE architectures. The motivation for development stems from the pain points of existing decision systems, such as high latency, lack of confidence metrics, and context limitations, aiming to create a decision model family that covers the full range from 0.8B to 35B, is deployable locally, and provides quantifiable outputs.
Core Value: Addresses the core issue in automated decision-making processes where "model outputs cannot be directly used for risk assessment." Each decision comes with an attached probability distribution, which can be used directly as confidence levels without additional calibration; at the same time, the mechanism of producing answers with a single forward pass compresses inference latency to the millisecond level, offering a new technical pathway for real-time decision-making scenarios.
Technical Features: Utilizes a linear attention architecture and the fast-linear-attention kernel, achieving extremely low latency for short requests; the 35B-A3B MoE model activates approximately 3B parameters per token through grouped matrix multiplication, resulting in latency roughly half that of the 27B dense model; probability fidelity quantization technology ensures that the GGUF quantized version matches the original weight decisions exactly on 99–100% of benchmark tasks.
2. Key Features
Typed Decision Output: Supports three structured question types: multiple-choice (choice), true/false (bool), and scoring (score). After a single forward inference, each question directly returns the probability distribution of all options, rather than a single answer text, providing downstream automated processes with directly quantifiable decision-making basis and eliminating the need for additional confidence calibration steps.
Multimodal State Input: State input supports not only text and JSON, but also images (updated on 2026-10-03). Evidence images and other visual data can be directly input into the model as part of the request, enabling the decision system to integrate visual information for comprehensive judgment. This significantly expands the model's applicability in scenarios such as document review, image recognition, and visual question answering.
Ultra-Long Context Handling: The entire series of models natively supports a prompt length of up to 262,144 tokens (256K). Long states can be read in one block without requiring multiple requests. This feature is particularly important for scenarios involving full contracts, long conversation histories, or large log files, avoiding information loss caused by truncating long texts.
Confidence Probability Output: The probability of each option is directly used as confidence level, making it suitable for automated processes requiring risk quantification. In scenarios such as financial risk control, medical decision support, and automated operations where the cost of errors is high, probability output allows the system to set thresholds for tiered processing, rather than making simple binary judgments.
High Throughput and Low Latency Inference: Short questions require only a few milliseconds for a single forward pass, with a 4B model latency of approximately 26ms and a 35B-A3B model latency of approximately 52.5ms. It supports batched inference and acceleration via CUDA Graphs, ensuring the server maintains stable low-latency performance in high-concurrency scenarios, making it ideal for high-throughput business integration.
Seamless Compatibility with Jev Ecosystem: Requests and responses use the TypeSafe
/v1/systemoneformat, allowing clients written for Jev to connect to the StartLux-Decision service without any modifications. This design significantly reduces migration costs, enabling existing Jev users to smoothly transition to an open-source, self-hostable decision-making solution.
3. How to Use
Environment Requirements: It is recommended to use a Linux server equipped with an NVIDIA GPU (CUDA support). Pre-installed Python 3.9+, PyTorch 2.1+, and CUDA 11.8 or higher are required. If using Apple Silicon devices, the model can be run via the MLX backend. Model weights can be downloaded from Hugging Face, ensuring sufficient disk space (approximately 8GB for the 4B model, and approximately 70GB for the 35B-A3B model).
Download Weights: Obtain the model weights via the command line from Hugging Face. Run
hf download startlux-models/StartLux-Decision-4B --local-dir StartLux-Decision-4Bto download the 4B model. To download other sizes, replace the model name with the appropriate one (e.g., StartLux-Decision-0.8B, StartLux-Decision-27B, etc.).Install Dependencies: Navigate to the model directory and execute
pip install -r requirements.txtto install the required dependency packages. Dependencies include fast kernel libraries such as flash-linear-attention and causal-conv1d, as well as web frameworks and serialization components needed for service operation.Verify Kernel Status: Run
python -m startlux_decision.check StartLux-Decision-4Bto verify the environment configuration. The output must include the text "fast kernels: active," indicating that the fast kernels have been correctly loaded. If they are not active, the server will refuse to start. This is a security mechanism designed to ensure inference performance.Start Inference Service: Execute
python -m startlux_decision.server --model StartLux-Decision-4B --port 8090to launch the local inference service. The service listens on port 8090 by default, but this can be customized using the--portparameter. Upon successful startup, the service will expose the/v1/systemoneendpoint for client calls.Send Decision Requests: Send a JSON request to
localhost:8090/v1/systemonecontainingstate(state description, which can be text, JSON, or an image URL) andquestions(a list of questions, supporting three types: choice, bool, and score). The response will return the probability distribution for each question's options, which can be directly used as decision confidence levels.Deployment with GGUF (Optional): If deployment is required in a llama.cpp environment, download the corresponding GGUF repository (quantization formats include BF16/Q8_0/Q4_K_M). First, start the
llama-server, then runpython -m startlux_decision.gguf_serverto provide decision services in front of llama.cpp. This method is suitable for resource-constrained environments or those already equipped with llama.cpp infrastructure.
4. Pros and Cons Analysis
| Pros |
|---|
| Probabilistic Decision Output: Directly returns the probability distribution of each option for every question, making confidence levels naturally usable without the need for additional calibration steps, providing structured output that can be directly consumed for risk quantification and automated decision-making. |
| Extremely Fast Inference Latency: Answers are generated with a single forward pass, with the 4B model requiring only 26ms and the 35B-A3B model only 52.5ms. Combined with batch inference and CUDA Graphs acceleration, it meets the strict latency requirements of real-time decision-making scenarios. |
| Support for Ultra-Long Context: The entire series natively supports input of up to 256K tokens. Long documents can be processed in one read, avoiding information loss caused by text truncation, making it suitable for scenarios such as contract review and log analysis. |
| Multimodal Decision-Making Capability: All model sizes can read images as a basis for decision-making, integrating visual and textual information for comprehensive judgment, offering a differentiated advantage in scenarios such as document review and image recognition. |
5. Comparative Analysis with Similar Tools
| Comparison Dimension | StartLux-Decision | Jev | SemIf |
|---|---|---|---|
| Product Positioning | Family of typed decision models (0.8B–35B, open source) | General-purpose decision system (closed API) | Open-source AI decision model, replicating Jev's semantic decision pattern |
| Decision Index 0.2.1 | 63.88 | 57.91 | To be announced |
| JevBench public (231 questions) | 208 questions passed | 199 questions passed | To be announced |
| Latency Performance | 3 questions per request: 102.3ms (H200 native) | 64.0ms (API gateway), end-to-end 109.7ms | To be announced |
| Context Length | 262,144 tokens (256K) | Not disclosed / shorter | To be announced |
| Image Input | Supported (full-size) | Not mentioned | To be announced |
| 38 Benchmark Comparisons | 31 benchmarks outperformed | 7 benchmarks outperformed | To be announced |
| Output Format | Probability distribution per option (confidence) | Option answer | Semantic decision |
| Weight Availability | HF + ModelScope, CC BY-NC 4.0 | Closed API | Open source |
| Deployment Method | Self-hosted + GGUF + MLX | Only cloud API | Self-hosted |
Selection Recommendations: For enterprise decision scenarios requiring local deployment, high data privacy requirements, and confidence output, StartLux-Decision is currently the standout choice in terms of overall capabilities. Its open-source weights, 256K context length, and multi-modal input capabilities provide clear advantages in scenarios such as financial risk control, enterprise search, and game AI. Particularly for real-time decision systems sensitive to latency, StartLux-Decision's millisecond-level inference and probabilistic output features can significantly simplify system architecture.
For users who are deeply integrated into the Jev ecosystem and have no concerns about data leaving the country, Jev's API gateway latency performance remains impressive. However, its closed nature and lack of confidence output limit its application in risk-sensitive scenarios. SemIf and APUS-OpenJev-v1, as replication and on-device solutions, are suitable for scenarios with special deployment location requirements or limited resources, but their benchmark performance and functional completeness still require validation through public data. Overall, StartLux-Decision currently holds a leading position in the open-source decision model space, though commercial licensing restrictions are key factors for enterprises to evaluate when implementing in production environments.
6. Editor's Summary
StartLux-Decision demonstrates clear technological innovation in the field of decision-making models. Its architecture design of "reading the answer in a single forward pass" breaks away from the traditional "generative response" paradigm, transforming decision-making problems into structured probability outputs. This design significantly reduces inference latency at the engineering level—4B model latency is 26ms, and 35B-A3B latency is 52.5ms. Combined with batch inference and CUDA Graphs acceleration, real-time decision-making becomes feasible. From the benchmark data, Decision Index 0.2.1 reaches 63.88, and JevBench public passes 208 out of 231 questions, outperforming 31 out of 38 comparisons. These figures support its leading position among open-source decision models.
In terms of practical value, the probabilistic output directly addresses the long-standing pain point of unavailable confidence levels in automated decision-making. The 256K context window and multimodal input further expand its application boundaries. With a family layout ranging from 0.8B to 35B, combined with GGUF quantization and MLX backend, it covers deployment needs from edge devices to the cloud, reflecting a detailed consideration of practical scenarios.
In terms of target users, this model is suitable for technical teams that need to build decision-making systems, especially developers in fields such as financial risk control, automated operations, game AI, and enterprise search. For teams already using Jev, the TypeSafe format compatibility design reduces migration costs. It is worth noting that the CC BY-NC 4.0 license imposes restrictions on commercial use, and public benchmark data for Chinese scenarios is still insufficient. Overall, StartLux-Decision sets a new technical benchmark for open-source decision models, and its future ecosystem development and commercial licensing strategies are worth continued attention.
7. Application Scenarios
Intelligent Customer Service Ticket Routing System: Input ticket text or screenshots, and a single request can determine the handling team, urgency, and severity. The output probabilities represent confidence levels, allowing the system to set thresholds for smooth transitions between automatic routing and manual intervention, significantly reducing customer service operational costs and improving response efficiency.
Computer Operation Automation: Drive a real browser to sequentially select controls and determine task completion status. The model has been demonstrated to automatically complete web-based shopping orders. Its multimodal capabilities enable it to understand page screenshots and DOM structures, with probability outputs used to determine "whether to continue the operation" or "if the task was successful," providing a decision-making core for RPA and agents.
Game AI Decision-Making: A single forward pass outputs probabilities for actions such as moving, building, and shooting, validated in real-time game scenarios like chess, StarCraft II, and Doom. Millisecond-level latency meets real-time gaming requirements, and probability outputs allow the AI to adjust strategies based on the uncertainty of the game situation, applicable to game testing, NPC intelligence enhancement, and more.
Financial Risk Control and Compliance: Covers contract clause verification, financial entity recognition, and phishing email detection. Confidence levels can be directly used in risk control processes, such as setting a 0.85 threshold to automatically intercept high-risk transactions, and routing lower-confidence cases to manual review. The 256K context window allows the model to fully process long contracts, while multimodal input supports document image verification.
Enterprise Search and Classification: Millisecond-level latency enables intent recognition, query classification, and RAG fact verification. In high-concurrency business scenarios, the model can perform real-time classification and routing of massive queries. Probability outputs can be used to measure the relevance of search results to the query, enhancing the accuracy of enterprise search and knowledge management.
8. FAQ
Q: What is the fundamental difference between StartLux-Decision and general-purpose large language models (e.g., GPT, Llama) in decision-making tasks?
A: General-purpose large models are centered around text generation, with outputs in natural language, making it difficult to directly quantify confidence levels. StartLux-Decision is optimized at the architectural level for decision-making tasks, rendering options as letter tokens and directly outputting the probability distribution of each option through a single forward inference. The output is in structured probability format, which can be used directly in automated decision-making processes without additional parsing or calibration.
Q: What input formats does the model support? How are images processed?
A: The model supports three input formats for state input: text, JSON, and images. Images are directly passed into the model as part of the request, alongside textual evidence, to serve as the basis for decision-making. All model sizes support image input, a feature enabled by a multimodal encoder that allows the model to integrate visual information for comprehensive judgment.
Q: How can we ensure that the model's decision-making capability is not compromised after quantization?
A: StartLux-Decision employs probability fidelity technology for quantization. The GGUF formats BF16 and Q8_0 produce identical decisions to the original weights on 99–100% of the JevBench project. The Q4_K_M quantization maintains decision-making capability while compressing the model's size, demonstrating high robustness of the model's decision-making ability to quantization. For critical scenarios, it is recommended to use BF16 or Q8_0 to ensure fully consistent decision outcomes.
Q: What advantages does the 35B-A3B Mixture-of-Experts (MoE) model have over the 27B dense model?
A: The 35B-A3B model activates only about 3B parameters per token, running the expert layers using grouped matrix multiplication (grouped matmul) and being covered by CUDA Graphs. This results in inference latency approximately half that of the 27B dense model, while maintaining higher model capacity and decision-making capability. It is an optimal configuration for latency-sensitive scenarios.
Q: Can the model be used for commercial purposes? Are there any license restrictions?
A: The model weights are licensed under CC BY-NC 4.0, which restricts their use to non-commercial purposes only. For commercial use in production environments, enterprises should contact StartLux Labs to obtain a commercial license. The license terms for the code repository are subject to the LICENSE file in the GitHub repository.
Q: How can the model be deployed on Apple Silicon devices? What is the performance like?
A: The model can be run on Apple Silicon devices using the MLX backend. MLX is a machine learning framework developed by Apple, optimized for Apple Silicon. The deployment method involves downloading the GGUF quantized model and loading it via the MLX backend. Performance depends on the chip model, with M-series Pro/Max chips capable of smoothly running medium and small-sized models.
Q: How can existing Jev clients be migrated to StartLux-Decision?
A: StartLux-Decision uses the TypeSafe /v1/systemone format for requests and responses, which is fully compatible with Jev. Simply change the client's API endpoint from the Jev cloud address to a local service address (e.g., localhost:8090/v1/systemone), and the client can be integrated without any modifications, requiring no changes to the client-side code.
9. Project Links
- GitHub Repository: https://github.com/StartLuxLabs/Startlux-Decision
- Hugging Face Model Library: https://huggingface.co/collections/startlux-models/startlux-decision-6abba92b301b573fa154d493
Related AI Model Articles

Mistral Large 4 – Mistral AI's Most Flagship Open-Source Large Model
Mistral Large 4 is the latest flagship open-source large model launched by Mistral AI, a French artificial intelligence company, and is nicknamed "Le Chonk" (the chubby cat). This model employs a fine...

IQuest-Q1 Review: A 320B Sparse MoE Open-Source Agent Foundation Model Specializing in Code and Long-Horizon Agent Tasks
IQuest-Q1 is an open-source Agent foundation model developed by IQuestLab, with a core focus on code generation and Agent task execution. The model employs a sparse MoE architecture, featuring a total...

In-Depth Review of M3.1-Flash-Preview: MiniMax's Text Programming Model for Everyday Development Scenarios
M3.1-Flash-Preview is the latest text programming model launched by MiniMax, initially released on the MiniMax Code intelligent programming client. Designed for everyday development scenarios, this mo...

R2T2 – In-Depth Review of NetEase Youdao's Open-Source Low-Latency True Streaming Speech Recognition Model
R2T2 (Real Real-Time Transcription) is a low-latency true streaming speech recognition model open-sourced by NetEase Youdao, built upon the Qwen3-ASR architecture. It employs an append-only mechanism ...
© All Rights Reserved. Some content on this site is partially generated by AI with human review.
