Back to Model List

Kev – Open-Source Decision Model Family in the Style of Jev, Supporting Self-Training and Self-Deployment

AI Tech Editorial
RSS Feed

Executive Summary:

Kev is an open-source family of decision models released by Jared Palmer, Vice President of Engineering at Cognition. It is built upon the Qwen3.5/Qwen3.8 base and offers multiple parameter scales ran...

1. What is Kev

Kev is an open-source family of decision models released by Jared Palmer, Vice President of Engineering at Cognition. It is built upon the Qwen3.5/Qwen3.8 base and offers multiple parameter scales ranging from 0.8B to 27B. Unlike traditional large language models, Kev does not generate natural language text. Instead, it receives state text and structured questions, and directly outputs calibrated probability distributions through a single forward pass. It supports three types of decision primitives: yes/no judgment, option routing, and rating scoring. Its API is compatible with TypeSafe System One. The core value of this project lies in transforming the "essay question" paradigm of large models into a "multiple-choice question" format, achieving high-concurrency decision scenarios with extremely low inference costs and millisecond-level latency. Additionally, the weights, training code, and evaluation datasets are fully open-sourced under the Apache-2.0 license.

Technical Positioning and Domain: Kev belongs to the domain of decision-specific models, offering a differentiated positioning from general-purpose conversational models. Its core innovation lies in combining a discriminative pointer head with a frozen Qwen base, enabling a structural transition from generative to discriminative architectures, and providing a high-cost-performance technical path for structured decision tasks.

Development Background: The project was led by Jared Palmer, VP of Engineering at Cognition, whose team has accumulated extensive experience in AI engineering and system design. The architecture of Kev is inspired by Jev (a closed-source decision model launched by TypeSafe), with its implementation ideas being reproduced and open-sourced. The goal is to break the monopoly of commercial APIs in decision-making scenarios. The official has disclosed the complete training cost, with the total cost for migrating and training three model sizes being approximately $95, and individual fine-tuning costing around $1, demonstrating a high level of cost transparency.

Core Value: Kev addresses three major pain points of traditional LLMs in decision-making scenarios: high latency and cost due to token-by-token generation, automation challenges caused by unreliable probability outputs, and data privacy and vendor lock-in issues from closed-source APIs. By outputting calibrated probabilities in a single forward pass, Kev compresses decision latency to the millisecond level and supports fully localized deployment, making high-throughput decision automation feasible.

Technical Features: Kev employs a lightweight architecture combining a frozen Qwen base, a rank-16 LoRA adapter, and a pointer head, resulting in extremely low training costs and reproducibility. The block causal attention mechanism enables parallel isolation of multiple questions, allowing any number of questions to be answered in a single forward pass. It also includes built-in temperature calibration and a frozen evaluation set, ensuring the reliability and audibility of the output probabilities.

2. Key Features

  • Structured Decision Output: Kev's core capability lies in directly outputting calibrated probabilities instead of generating text. The system receives a state description along with several structured questions, and a single forward pass can return the probability distribution for each candidate option, completely bypassing the token-by-token generation process. This design reduces decision latency from seconds to milliseconds, with only 18ms required to answer six questions on the H100, while also eliminating the common "nonsense" issues in generative models, making the output directly usable in automated decision pipelines.

  • Three Types of Decision Primitives: Kev natively supports three decision types: noul (yes/no judgment), choice (multi-option routing), and score (rating scoring). noul is suitable for binary classification scenarios such as content compliance checks, choice is used for multi-category routing like ticket classification, and score is used for continuous dimension evaluation such as risk rating scoring. These three primitives can be freely combined to cover the majority of structured decision-making needs.

  • Batch Question Isolation Mechanism: Multiple questions can share the same input state, while being isolated from each other through block causal attention masks. Each token can only read state information and content within its own question block, and position IDs are reset for each question, ensuring that all questions are answered in a single forward pass without interfering with each other. On architectures like Qwen3.5/3.8 that use a hybrid Gated DeltaNet design, the system instead processes each question independently in its own row while reusing the state KV cache, providing stricter isolation and significantly improving multi-task parallel efficiency.

  • Local Deployment Capabilities: The system provides the kev.serve command to launch the service with a single line, exposing the POST /v1/systemone endpoint. It is compatible with three hardware platforms: CUDA, ROCm, and Apple Silicon (MLX). Users can run the model on their own servers, ensuring data never leaves the premises, meeting the privacy requirements of high-compliance scenarios such as finance and government affairs. It also supports one-click deployment of HTTPS endpoints on cloud platforms like Modal, automatically scaling down to zero when idle and not incurring any charges.

  • Low-Cost Fine-Tuning and Agent Automation: The system supports continuing training using proprietary labeled data based on official checkpoints. The accompanying kev-finetune skill enables coding agents to automatically complete the entire process of data preparation, training, evaluation, and deployment. Users do not need local GPUs to trigger automated fine-tuning via npx skills add jaredpalmer/kev@kev-finetune. Official data shows that the cost of a single fine-tuning session is approximately $1, greatly reducing the entry barrier for customized decision models.

  • Built-in Evaluation and Calibration System: Each checkpoint fits temperature parameters on held-out data, and during inference, directly outputs calibrated probabilities. Temperature adjustment does not affect the ranking but only calibrates the confidence scores. The project includes a frozen evaluation set, reporting four metrics: accuracy, Brier score, calibration error, and the proportion of decisions that can be automated. The CI process automatically verifies and publishes these metrics, ensuring performance is traceable and comparable during model iterations.

3. How to Use

  1. Environment Setup: Requires a Python 3.12 or 3.13 environment and the installation of the uv package manager. After cloning the GitHub repository, run uv sync --extra serve to install all dependencies. In terms of hardware, the 0.8B model can run smoothly on Apple Silicon, while models of 4B and above are recommended to be used with an NVIDIA GPU (CUDA) or AMD GPU (ROCm). The required GPU memory varies from 2GB to 24GB depending on the model size.

  2. Launch Local Service: Run KEV_DTYPE=bf16 uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b --port 8009 to start the inference service. On the first run, the model weights will be automatically downloaded from Hugging Face. The --run parameter specifies the model identifier, and --port specifies the service port. Once the service is successfully launched, it listens on the /v1/systemone endpoint, waiting to receive decision requests.

  3. Send Decision Requests: Submit a JSON-formatted request body to POST /v1/systemone, which includes the state field (state text) and the questions field (array of structured questions). The service returns the calibrated probability distribution for each option. For example, input a customer complaint text along with three questions: "Should we escalate to a human?", "What type of complaint is this?", and "What is the severity level?"—a single request will provide the probability distribution for all answers.

  4. Python SDK Usage: Use the official TypeSafe SDK, and set the base_url to point to the local service address (e.g., http://localhost:8009). This allows existing Jev application code to be reused with zero modifications for switching. The SDK automatically handles request serialization and response parsing, so developers don't need to worry about the details of the underlying HTTP protocol.

  5. Browser Demo Experience: Access the Hugging Face Space for an online demo experience, or run the local playground script to input state text and questions through a graphical interface and visually inspect the probability output. This is ideal for quickly verifying model performance or conducting technical evaluations.

  6. Fine-tuning with Custom Data: Organize annotated data into JSONL format (each line includes the state text, questions, and the correct option label), then execute uv run python -m kev.train --init_from jaredpalmer/kev-4b ... to continue training from the official checkpoint. During training, the base model weights are frozen, and only the LoRA adapter and pointer head are updated, resulting in low memory usage. The cost of a single fine-tuning session is approximately $1.

  7. Agent-based Automatic Fine-tuning: Run npx skills add jaredpalmer/kev@kev-finetune to install the fine-tuning skill, and then let the coding agent automatically complete the entire process of question design, data generation, model training, evaluation, and deployment. This mode does not require a local GPU and is suitable for teams without deep learning expertise who need a customized decision-making model.

  8. Cloud Deployment and Evaluation: Execute pip install modal && modal deploy kev_serve.py to deploy the service as an HTTPS endpoint, which supports automatic scaling down to zero. Before going live, it is recommended to run uv run python -m kev.benchmark --run <model> --suite evals/v4/transfer-v4 on the frozen evaluation dataset to verify accuracy, Brier score, and calibration error. Ensure the model quality meets the standards before deploying it into production.

4. Pros and Cons Analysis

Pros
Extremely low inference cost: A single forward pass directly outputs probabilities, completely skipping token-by-token generation. On the H100, answering six questions takes only 18ms. After self-hosting, the invocation cost approaches zero, offering a clear advantage over the API's pay-per-use model.
Transparent and reproducible training cost: The official provides a complete training cost breakdown. Migrating and training across three model tiers costs approximately $95, while individual fine-tuning costs about $1 per session. Under the Apache-2.0 license, all training code and weights are open-sourced, supporting full reproducibility and independent iteration.
Data privacy and autonomy: Full local deployment is supported, ensuring data stays within the organization and avoiding vendor lock-in risks. It is compatible with the TypeSafe System One API, allowing existing Jev applications to switch without modification, resulting in low migration costs.
High efficiency for parallel multi-question processing: The block causal attention mechanism enables shared state computation across multiple questions without cross-contamination. A single forward pass can answer any number of questions, significantly improving batch decision-making throughput.
Friendly open-source license: The Apache-2.0 license permits commercial use and free modification. All weights, code, and evaluation datasets are publicly available, offering strong auditability and making it suitable for enterprises with strict compliance requirements.

5. Comparative Analysis with Similar Tools

Comparison Dimension Kev (Open Source) Jev (Closed-Source Commercial API) General LLM (e.g., GPT-4o)
Core Architecture Qwen3.5/3.8 base + LoRA + pointer head, discriminative probability output Not disclosed, Kev replicates its architectural ideas Dense Transformer, autoregressive generation
Model Tier Four tiers available: 0.8B / 4B / 9B / 27B Single hosted model Multiple tiers of general models
Decision Accuracy 27B tier: 0.848, 4B tier: 0.838, 9B tier: 0.852 (new source evaluation set) Development set: 0.857 Task-dependent, requires prompt engineering and post-processing
Calibration Quality Confidence error rate: 4.0%, automation rate: 0.45~0.57 under 5% error budget Confidence error rate: 3.7%, automation rate: 0.70 Calibration unstable, requires additional temperature scaling or Platt scaling
Inference Speed 6 questions in 18ms on H100, supports batch inference Limited by API network latency, usually hundreds of milliseconds Token-by-token generation, second-level latency
Deployment Method Local / On-premise server / Modal cloud, supports CUDA, ROCm, MLX Cloud API, data passes through third-party Cloud API or local deployment (requires large GPU memory)
Single Call Cost Weights are free, zero call cost with self-hosting Pay-as-you-go based on usage Pay-per-token
Data Privacy Fully localized, data does not leave the premises Data passes through third-party API Data passes through third-party API
Open Source License Apache-2.0, full open weights and code Closed-source commercial license Partially open, commercial use requires license
Community Ecosystem Early stage, GitHub open, community growing rapidly Commercial support, mature ecosystem Large ecosystem, rich toolchain

Selection Recommendations: For enterprises prioritizing data privacy, cost control, and high-throughput decision-making, Kev is currently the preferred open-source solution. Its API compatibility with TypeSafe System One allows smooth migration for existing Jev users. If the team lacks deep learning infrastructure, managed APIs like Jev can be quickly integrated, but at the cost of data leakage and ongoing usage expenses. For decision-making scenarios requiring interpretability or complex reasoning, a general LLM with structured output constraints remains a necessary choice. Traditional classification models maintain cost-performance advantages in scenarios with clear features and sufficient data volume, but cannot handle unstructured text inputs.

6. Editor's Summary

Kev's release provides a noteworthy open-source technical path for decision-making AI applications. From a technological innovation perspective, the combination of the pointer head and frozen base breaks away from the conventional mindset of "fine-tuning generative models." By restructuring the architecture, it transforms the generative problem into a discriminative one, achieving a two-order-of-magnitude increase in inference speed while maintaining semantic understanding capabilities. The multi-problem isolation design of the block causal attention mechanism also demonstrates thoughtful engineering, fully utilizing computational efficiency for batch decision-making. The transparency of training costs is a positive example for the open-source community—only $95 for model porting training and $1 per fine-tuning session, bringing the customization threshold of decision models down to a level that individual developers can afford.

From a practical value standpoint, Kev directly addresses three rigid requirements in production environments: millisecond-level latency to support high-concurrency scenarios, calibrated probabilities to ensure the reliability of automated decision-making, and local deployment to meet data compliance requirements. Its API compatibility with TypeSafe System One reduces migration costs, allowing existing Jev users to quickly switch to a self-hosted solution. The built-in frozen validation set and CI verification mechanism also reflect an engineering-oriented mindset, providing a quantifiable quality benchmark for model iteration.

In terms of target users, Kev is particularly suitable for three types of teams: first, financial and government institutions under data privacy compliance pressure; second, customer service and review platforms requiring high-throughput decision-making capabilities; and third, researchers and independent developers looking to conduct low-cost experiments in the decision model domain. For teams seeking out-of-the-box usability, the hosted API remains a more convenient choice, but Kev offers an alternative path with full autonomy and control.

Looking ahead, Kev's development potential hinges on three directions: whether continuous upgrades to the base model can lead to accuracy breakthroughs, whether the community can accumulate a rich set of industry-specific fine-tuning datasets, and whether multi-modal decision support will be included in the roadmap. The current design, ranging from 0.8B to 27B parameters, already covers deployment needs from edge devices to cloud clusters. If Kev continues to iterate on Chinese scenario optimization and long-text state understanding, it has the potential to become a critical infrastructure in the decision model field.

7. Application Scenarios

  • Intelligent Ticket Routing for Customer Service: A single forward pass simultaneously determines three questions: "Which department to route to," "Whether to escalate to a human agent," and "Customer anger level." Tickets are automatically routed based on probability thresholds. Tickets with low confidence are automatically forwarded to human agents, while those with high confidence are fully automated, significantly reducing the workload of the customer service team. After integration with an e-commerce platform, the average processing time for tickets was reduced from minutes to seconds.

  • UGC Content Safety Review: Performs three types of judgments in parallel on user-generated content: "Whether it is non-compliant" (noul), "Type of non-compliance" (choice), and "Severity level" (score). Content with high confidence of being non-compliant is automatically taken down, while content with medium or low confidence enters a human review queue. The probability distribution helps reviewers quickly identify issues. Compared to keyword filtering solutions, it can effectively detect semantic variations and subtle expressions.

  • Preliminary Review for Financial Credit: In credit application processing, identifies the type of complaint, whether human intervention is needed, and the risk level. The probability distribution is not only used for automated decision-making but also serves as a compliance audit trail, assisting in determining the priority for human review. Banks can fully deploy the system locally, ensuring customer data remains within the domain, meeting financial regulatory requirements for data security.

  • E-commerce Intent Classification and Multi-Label Routing: Routes user queries to downstream systems such as search, recommendation, or after-sales service, effectively handling multi-label distribution issues where a single term may fall into multiple categories. For example, a query about "returning goods" may simultaneously involve after-sales policies, logistics tracking, and refund status. Kev outputs the probability distribution for each intent, allowing downstream systems to make flexible decisions and support more nuanced interaction strategies.

8. FAQ

Q: What is the core difference between Kev and general-purpose large language models?
A: General-purpose LLMs generate text token by token using autoregressive methods, producing natural language responses. Kev completely skips the generation process and instead scores candidate options using a pointer head and outputs calibrated probabilities. This discriminative design makes Kev faster (millisecond-level), more cost-effective, and provides more reliable probabilities for decision-making tasks, but at the cost of losing text generation and free conversation capabilities.

Q: How does Kev's accuracy compare to Jev's?
A: According to officially published evaluation data, Kev 27B has an accuracy of 0.848, Kev 9B has an accuracy of 0.852, and Kev 4B has an accuracy of 0.838 on the new source evaluation dataset. Jev has an accuracy of 0.857 on the development set. Both are at similar levels, but Kev's evaluation dataset is a frozen public dataset, allowing reproducible verification of results, whereas Jev's evaluation methodology is not fully disclosed.

Q: How much data is needed to fine-tune Kev? What is the cost?
A: The official recommendation is to start with hundreds of annotated samples, in JSONL format (status text + question + correct label). The training process freezes the base model and only updates the LoRA adapter and the pointer head. The cost of a single fine-tuning session is approximately $1 (based on cloud GPU billing). The actual data volume required depends on the task complexity; for complex decision-making scenarios, it is recommended to accumulate several thousand high-quality annotated samples.

Q: Which hardware platforms does Kev support?
A: Kev supports CUDA (NVIDIA GPU), ROCm (AMD GPU), and Apple Silicon (with acceleration via the MLX framework). The 0.8B model can run smoothly on Apple Silicon, while models of 4B and above are recommended to be used with GPUs having at least 8GB of VRAM. CPU inference is possible but slower, and is not suitable for production environments.

Q: How can I migrate existing Jev applications to Kev?
A: Kev's API is fully compatible with TypeSafe System One. Applications using the official SDK only need to point the base_url to the local Kev service address, with no code changes required. Kev's response format is consistent with Jev, including calibrated probabilities for each option. Therefore, existing post-processing logic and decision-making workflows can be directly reused.

Q: Can Kev's output probabilities be used directly for automated decision-making?
A: Yes. Kev includes an internal temperature calibration mechanism, and its output probabilities are calibrated, making the confidence scores meaningful in practice. The official documentation provides a "proportion of automatable decisions" metric, which indicates the percentage of samples that can be safely automated under a given error budget. It is recommended to set a probability threshold based on your business tolerance for errors, and route samples below that threshold to manual processing.

9. Project Links

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.