Back to Model List

In-Depth Review of Startlux-Decision – StartLux Labs' Open-Source Decision Model Family

AI Tech Editorial
RSS Feed

Executive Summary:

Startlux-Decision is an open-source decision model family introduced by StartLux Labs, covering five dense models ranging from 0.8B to 27B, as well as a 35B-A3B mixed expert model. It is specifically ...

1. What is Startlux-Decision

Startlux-Decision is an open-source decision model family introduced by StartLux Labs, covering five dense models ranging from 0.8B to 27B, as well as a 35B-A3B mixed expert model. It is specifically designed for typed decision-making tasks. This model family accepts state inputs in the form of text, JSON, or images, and directly returns the probability distribution of each option through a single forward inference for structured questions such as selection, yes/no, and scoring. This makes it naturally suitable for use as confidence levels. The model natively supports ultra-long context of up to 256K tokens, has millisecond-level inference latency, and is compatible with Jev's TypeSafe client format, providing an efficient and quantifiable technical foundation for automated decision-making and agent applications.

Technical Positioning and Domain: Belongs to the intersection of natural language processing and decision intelligence, focusing on shifting the capabilities of large language models from open-ended generation to structured, quantifiable decision outputs. Unlike general-purpose conversational models, Startlux-Decision is optimized at the architectural level for the "read state – output probability" paradigm, making it suitable for automated scenarios requiring determinism, explainability, and risk quantification.

Development Background: Developed by the StartLux Labs team, which has deep expertise in efficient inference, linear attention mechanisms, and MoE architectures. The motivation for development stems from the pain points of existing decision systems, such as high latency, lack of confidence metrics, and context limitations, aiming to create a decision model family that covers the full range from 0.8B to 35B, is deployable locally, and provides quantifiable outputs.

Core Value: Addresses the core issue in automated decision-making processes where "model outputs cannot be directly used for risk assessment." Each decision comes with an attached probability distribution, which can be used directly as confidence levels without additional calibration; at the same time, the mechanism of producing answers with a single forward pass compresses inference latency to the millisecond level, offering a new technical pathway for real-time decision-making scenarios.

Technical Features: Utilizes a linear attention architecture and the fast-linear-attention kernel, achieving extremely low latency for short requests; the 35B-A3B MoE model activates approximately 3B parameters per token through grouped matrix multiplication, resulting in latency roughly half that of the 27B dense model; probability fidelity quantization technology ensures that the GGUF quantized version matches the original weight decisions exactly on 99–100% of benchmark tasks.

2. Key Features

  • Typed Decision Output: Supports three structured question types: multiple-choice (choice), true/false (bool), and scoring (score). After a single forward inference, each question directly returns the probability distribution of all options, rather than a single answer text, providing downstream automated processes with directly quantifiable decision-making basis and eliminating the need for additional confidence calibration steps.

  • Multimodal State Input: State input supports not only text and JSON, but also images (updated on 2026-10-03). Evidence images and other visual data can be directly input into the model as part of the request, enabling the decision system to integrate visual information for comprehensive judgment. This significantly expands the model's applicability in scenarios such as document review, image recognition, and visual question answering.

  • Ultra-Long Context Handling: The entire series of models natively supports a prompt length of up to 262,144 tokens (256K). Long states can be read in one block without requiring multiple requests. This feature is particularly important for scenarios involving full contracts, long conversation histories, or large log files, avoiding information loss caused by truncating long texts.

  • Confidence Probability Output: The probability of each option is directly used as confidence level, making it suitable for automated processes requiring risk quantification. In scenarios such as financial risk control, medical decision support, and automated operations where the cost of errors is high, probability output allows the system to set thresholds for tiered processing, rather than making simple binary judgments.

  • High Throughput and Low Latency Inference: Short questions require only a few milliseconds for a single forward pass, with a 4B model latency of approximately 26ms and a 35B-A3B model latency of approximately 52.5ms. It supports batched inference and acceleration via CUDA Graphs, ensuring the server maintains stable low-latency performance in high-concurrency scenarios, making it ideal for high-throughput business integration.

  • Seamless Compatibility with Jev Ecosystem: Requests and responses use the TypeSafe /v1/systemone format, allowing clients written for Jev to connect to the StartLux-Decision service without any modifications. This design significantly reduces migration costs, enabling existing Jev users to smoothly transition to an open-source, self-hostable decision-making solution.

3. How to Use

  1. Environment Requirements: It is recommended to use a Linux server equipped with an NVIDIA GPU (CUDA support). Pre-installed Python 3.9+, PyTorch 2.1+, and CUDA 11.8 or higher are required. If using Apple Silicon devices, the model can be run via the MLX backend. Model weights can be downloaded from Hugging Face, ensuring sufficient disk space (approximately 8GB for the 4B model, and approximately 70GB for the 35B-A3B model).

  2. Download Weights: Obtain the model weights via the command line from Hugging Face. Run hf download startlux-models/StartLux-Decision-4B --local-dir StartLux-Decision-4B to download the 4B model. To download other sizes, replace the model name with the appropriate one (e.g., StartLux-Decision-0.8B, StartLux-Decision-27B, etc.).

  3. Install Dependencies: Navigate to the model directory and execute pip install -r requirements.txt to install the required dependency packages. Dependencies include fast kernel libraries such as flash-linear-attention and causal-conv1d, as well as web frameworks and serialization components needed for service operation.

  4. Verify Kernel Status: Run python -m startlux_decision.check StartLux-Decision-4B to verify the environment configuration. The output must include the text "fast kernels: active," indicating that the fast kernels have been correctly loaded. If they are not active, the server will refuse to start. This is a security mechanism designed to ensure inference performance.

  5. Start Inference Service: Execute python -m startlux_decision.server --model StartLux-Decision-4B --port 8090 to launch the local inference service. The service listens on port 8090 by default, but this can be customized using the --port parameter. Upon successful startup, the service will expose the /v1/systemone endpoint for client calls.

  6. Send Decision Requests: Send a JSON request to localhost:8090/v1/systemone containing state (state description, which can be text, JSON, or an image URL) and questions (a list of questions, supporting three types: choice, bool, and score). The response will return the probability distribution for each question's options, which can be directly used as decision confidence levels.

  7. Deployment with GGUF (Optional): If deployment is required in a llama.cpp environment, download the corresponding GGUF repository (quantization formats include BF16/Q8_0/Q4_K_M). First, start the llama-server, then run python -m startlux_decision.gguf_server to provide decision services in front of llama.cpp. This method is suitable for resource-constrained environments or those already equipped with llama.cpp infrastructure.

4. Pros and Cons Analysis

Pros
Probabilistic Decision Output: Directly returns the probability distribution of each option for every question, making confidence levels naturally usable without the need for additional calibration steps, providing structured output that can be directly consumed for risk quantification and automated decision-making.
Extremely Fast Inference Latency: Answers are generated with a single forward pass, with the 4B model requiring only 26ms and the 35B-A3B model only 52.5ms. Combined with batch inference and CUDA Graphs acceleration, it meets the strict latency requirements of real-time decision-making scenarios.
Support for Ultra-Long Context: The entire series natively supports input of up to 256K tokens. Long documents can be processed in one read, avoiding information loss caused by text truncation, making it suitable for scenarios such as contract review and log analysis.
Multimodal Decision-Making Capability: All model sizes can read images as a basis for decision-making, integrating visual and textual information for comprehensive judgment, offering a differentiated advantage in scenarios such as document review and image recognition.

5. Comparative Analysis with Similar Tools

Comparison Dimension StartLux-Decision Jev SemIf
Product Positioning Family of typed decision models (0.8B–35B, open source) General-purpose decision system (closed API) Open-source AI decision model, replicating Jev's semantic decision pattern
Decision Index 0.2.1 63.88 57.91 To be announced
JevBench public (231 questions) 208 questions passed 199 questions passed To be announced
Latency Performance 3 questions per request: 102.3ms (H200 native) 64.0ms (API gateway), end-to-end 109.7ms To be announced
Context Length 262,144 tokens (256K) Not disclosed / shorter To be announced
Image Input Supported (full-size) Not mentioned To be announced
38 Benchmark Comparisons 31 benchmarks outperformed 7 benchmarks outperformed To be announced
Output Format Probability distribution per option (confidence) Option answer Semantic decision
Weight Availability HF + ModelScope, CC BY-NC 4.0 Closed API Open source
Deployment Method Self-hosted + GGUF + MLX Only cloud API Self-hosted

Selection Recommendations: For enterprise decision scenarios requiring local deployment, high data privacy requirements, and confidence output, StartLux-Decision is currently the standout choice in terms of overall capabilities. Its open-source weights, 256K context length, and multi-modal input capabilities provide clear advantages in scenarios such as financial risk control, enterprise search, and game AI. Particularly for real-time decision systems sensitive to latency, StartLux-Decision's millisecond-level inference and probabilistic output features can significantly simplify system architecture.

For users who are deeply integrated into the Jev ecosystem and have no concerns about data leaving the country, Jev's API gateway latency performance remains impressive. However, its closed nature and lack of confidence output limit its application in risk-sensitive scenarios. SemIf and APUS-OpenJev-v1, as replication and on-device solutions, are suitable for scenarios with special deployment location requirements or limited resources, but their benchmark performance and functional completeness still require validation through public data. Overall, StartLux-Decision currently holds a leading position in the open-source decision model space, though commercial licensing restrictions are key factors for enterprises to evaluate when implementing in production environments.

6. Editor's Summary

StartLux-Decision demonstrates clear technological innovation in the field of decision-making models. Its architecture design of "reading the answer in a single forward pass" breaks away from the traditional "generative response" paradigm, transforming decision-making problems into structured probability outputs. This design significantly reduces inference latency at the engineering level—4B model latency is 26ms, and 35B-A3B latency is 52.5ms. Combined with batch inference and CUDA Graphs acceleration, real-time decision-making becomes feasible. From the benchmark data, Decision Index 0.2.1 reaches 63.88, and JevBench public passes 208 out of 231 questions, outperforming 31 out of 38 comparisons. These figures support its leading position among open-source decision models.

In terms of practical value, the probabilistic output directly addresses the long-standing pain point of unavailable confidence levels in automated decision-making. The 256K context window and multimodal input further expand its application boundaries. With a family layout ranging from 0.8B to 35B, combined with GGUF quantization and MLX backend, it covers deployment needs from edge devices to the cloud, reflecting a detailed consideration of practical scenarios.

In terms of target users, this model is suitable for technical teams that need to build decision-making systems, especially developers in fields such as financial risk control, automated operations, game AI, and enterprise search. For teams already using Jev, the TypeSafe format compatibility design reduces migration costs. It is worth noting that the CC BY-NC 4.0 license imposes restrictions on commercial use, and public benchmark data for Chinese scenarios is still insufficient. Overall, StartLux-Decision sets a new technical benchmark for open-source decision models, and its future ecosystem development and commercial licensing strategies are worth continued attention.

7. Application Scenarios

  • Intelligent Customer Service Ticket Routing System: Input ticket text or screenshots, and a single request can determine the handling team, urgency, and severity. The output probabilities represent confidence levels, allowing the system to set thresholds for smooth transitions between automatic routing and manual intervention, significantly reducing customer service operational costs and improving response efficiency.

  • Computer Operation Automation: Drive a real browser to sequentially select controls and determine task completion status. The model has been demonstrated to automatically complete web-based shopping orders. Its multimodal capabilities enable it to understand page screenshots and DOM structures, with probability outputs used to determine "whether to continue the operation" or "if the task was successful," providing a decision-making core for RPA and agents.

  • Game AI Decision-Making: A single forward pass outputs probabilities for actions such as moving, building, and shooting, validated in real-time game scenarios like chess, StarCraft II, and Doom. Millisecond-level latency meets real-time gaming requirements, and probability outputs allow the AI to adjust strategies based on the uncertainty of the game situation, applicable to game testing, NPC intelligence enhancement, and more.

  • Financial Risk Control and Compliance: Covers contract clause verification, financial entity recognition, and phishing email detection. Confidence levels can be directly used in risk control processes, such as setting a 0.85 threshold to automatically intercept high-risk transactions, and routing lower-confidence cases to manual review. The 256K context window allows the model to fully process long contracts, while multimodal input supports document image verification.

  • Enterprise Search and Classification: Millisecond-level latency enables intent recognition, query classification, and RAG fact verification. In high-concurrency business scenarios, the model can perform real-time classification and routing of massive queries. Probability outputs can be used to measure the relevance of search results to the query, enhancing the accuracy of enterprise search and knowledge management.

8. FAQ

Q: What is the fundamental difference between StartLux-Decision and general-purpose large language models (e.g., GPT, Llama) in decision-making tasks?
A: General-purpose large models are centered around text generation, with outputs in natural language, making it difficult to directly quantify confidence levels. StartLux-Decision is optimized at the architectural level for decision-making tasks, rendering options as letter tokens and directly outputting the probability distribution of each option through a single forward inference. The output is in structured probability format, which can be used directly in automated decision-making processes without additional parsing or calibration.

Q: What input formats does the model support? How are images processed?
A: The model supports three input formats for state input: text, JSON, and images. Images are directly passed into the model as part of the request, alongside textual evidence, to serve as the basis for decision-making. All model sizes support image input, a feature enabled by a multimodal encoder that allows the model to integrate visual information for comprehensive judgment.

Q: How can we ensure that the model's decision-making capability is not compromised after quantization?
A: StartLux-Decision employs probability fidelity technology for quantization. The GGUF formats BF16 and Q8_0 produce identical decisions to the original weights on 99–100% of the JevBench project. The Q4_K_M quantization maintains decision-making capability while compressing the model's size, demonstrating high robustness of the model's decision-making ability to quantization. For critical scenarios, it is recommended to use BF16 or Q8_0 to ensure fully consistent decision outcomes.

Q: What advantages does the 35B-A3B Mixture-of-Experts (MoE) model have over the 27B dense model?
A: The 35B-A3B model activates only about 3B parameters per token, running the expert layers using grouped matrix multiplication (grouped matmul) and being covered by CUDA Graphs. This results in inference latency approximately half that of the 27B dense model, while maintaining higher model capacity and decision-making capability. It is an optimal configuration for latency-sensitive scenarios.

Q: Can the model be used for commercial purposes? Are there any license restrictions?
A: The model weights are licensed under CC BY-NC 4.0, which restricts their use to non-commercial purposes only. For commercial use in production environments, enterprises should contact StartLux Labs to obtain a commercial license. The license terms for the code repository are subject to the LICENSE file in the GitHub repository.

Q: How can the model be deployed on Apple Silicon devices? What is the performance like?
A: The model can be run on Apple Silicon devices using the MLX backend. MLX is a machine learning framework developed by Apple, optimized for Apple Silicon. The deployment method involves downloading the GGUF quantized model and loading it via the MLX backend. Performance depends on the chip model, with M-series Pro/Max chips capable of smoothly running medium and small-sized models.

Q: How can existing Jev clients be migrated to StartLux-Decision?
A: StartLux-Decision uses the TypeSafe /v1/systemone format for requests and responses, which is fully compatible with Jev. Simply change the client's API endpoint from the Jev cloud address to a local service address (e.g., localhost:8090/v1/systemone), and the client can be integrated without any modifications, requiring no changes to the client-side code.

9. Project Links

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.