Instructions to use Kronumos/Kronumos-Kairos-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Kronumos/Kronumos-Kairos-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Kronumos/Kronumos-Kairos-v2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Kronumos/Kronumos-Kairos-v2") model = AutoModelForCausalLM.from_pretrained("Kronumos/Kronumos-Kairos-v2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Kronumos/Kronumos-Kairos-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Kronumos/Kronumos-Kairos-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kronumos/Kronumos-Kairos-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Kronumos/Kronumos-Kairos-v2
- SGLang
How to use Kronumos/Kronumos-Kairos-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Kronumos/Kronumos-Kairos-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kronumos/Kronumos-Kairos-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Kronumos/Kronumos-Kairos-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Kronumos/Kronumos-Kairos-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use Kronumos/Kronumos-Kairos-v2 with Docker Model Runner:
docker model run hf.co/Kronumos/Kronumos-Kairos-v2
- ⚡ Kronumos 2 Kairos: The Dual-Brain Sub-Cortex Autonomous Program Repair Engine
⚡ Kronumos 2 Kairos: The Dual-Brain Sub-Cortex Autonomous Program Repair Engine
Organization: Tokenectomy Labs
Base Model: Qwen/Qwen2.5-Coder-7B-Instruct
Research Preprint: Springer Nature Research Square (DOI: 10.21203/rs.3.rs-11205335/v1)
Lead Author: Muhammad Naufal Daffa (ORCID: 0009-0000-7909-4916)
GitHub Repository: Tokenectomy-Labs/Kronomus
Kronumos 2 Kairos is an open-weight 7B cybernetic autonomous program repair (APR) model fine-tuned for high-precision code remediation on real-world production software bugs. It pairs parametric neural intuition with a deterministic, zero-allocation Rust Sub-Cortex (libtokenectomy_subcortex.so, C-ABI 5µs latency).
📣 Announcement: Right Brain, Procedural Seed Cortex (In Development)
Status: In active development. Not yet part of the released weights.
The Kronumos Dual-Brain architecture is getting its next major component: the Right Brain, a cortex built entirely on Procedural Seeds.
Today, Kronumos Kairos pairs a neural Cortex (Qwen2.5-Coder-7B reasoning) with a deterministic Sub-Cortex (native Rust AST slicing and healing). The Right Brain adds a third pillar that works on a different principle: instead of generating a repair from learned weights, it derives repair strategies from compact procedural seeds (<64 bytes each), expanded deterministically at runtime with no database lookups.
What this adds
| Component | Role | Nature |
|---|---|---|
| Left Brain (Cortex) | Semantic reasoning, 5-step CoT, code hunk generation | Parametric / neural |
| Right Brain (new) | Seed-driven strategy generation and pattern intuition | Procedural / seed-based |
| Sub-Cortex | AST validation, bracket and indent healing, Merkle ledger | Deterministic / Rust |
Design goals
- Zero-DB, seed-native: strategies are regenerated from seeds, not retrieved from storage.
- Reproducible: the same seed always yields the same strategy.
- Tiny footprint: seeds stay in the L1 cache, with no extra VRAM requirement.
- Integrated with the Dual-Key Consensus Gate: Right Brain proposals must pass the same semantic and AST validation as every other candidate.
Roadmap
- Procedural Cognitive Kernel (seed-based invariant diagnosis)
- Right Brain seed expander
- Integration into the Consensus Gate
- SWE-bench Verified re-evaluation with the Right Brain enabled
- Public release in the Kronumos Kairos family
Results and benchmarks for the Right Brain will be published only after full evaluation. Follow the GitHub repository for updates.
🏛️ The Kronumos Dual-Brain Family Portfolio
| Tier | Model | Parameters | Target Workload | Latency / Footprint |
|---|---|---|---|---|
| Edge / Workstation | Kronumos 2 Kairos | 7.6B | Rapid local bug remediation, offline laptops, CI/CD gates | <2.5s / 4-bit 5.5 GB VRAM |
| Edge / Sovereign | Kronumos 14B Kairos | 14.7B | Complex algebraic, cross-module AST repairs (sympy, sphinx) |
<4.5s / 4-bit 9.2 GB VRAM |
| Titan / Enterprise | Kronumos Aion | 671B MoE | Deep multi-hop counterfactual reasoning & frontier SWE-bench | Enterprise Cluster / Cloud API |
🥊 Benchmark Verification: SWE-bench Verified (500 Instances)
Evaluated end-to-end on the official Princeton SWE-bench Verified benchmark (500 production instances across Django, Scikit-Learn, PyData Xarray, Sphinx, Sympy, etc.) using official Docker execution containers.
| Metric | Kronumos 2 Kairos | Industry Multi-Turn Baselines |
|---|---|---|
| Model Size | 7B Parameters | 70B - 405B / Frontier APIs |
| Execution Mode | Single-Pass Zero-Shot | Multi-Turn Agent Loop (50-100 Turns) |
| Avg Tokens / Task | 2,512 Tokens | 40,000 - 150,000 Tokens |
| Token Efficiency | 93.5% Reduction | Baseline (1.0x) |
| API Cost | $0.00 (Pure Local Weights) | $3.00 - $15.00 per issue |
| Verified Resolved Tasks | 8 Full Production Issues | - |
🏆 Verified Resolved Production Issues:
django__django-13569: Broken aggregation expression logic in database queries.django__django-13658: Management command argument parser collision.django__django-14855: Admin URL generation prefix regression.django__django-15104: Model custom key migration constraint hazard.django__django-16333: Many-to-many relationship foreign key mapping.pydata__xarray-4629: Multi-index coordinate slice dimension regression.scikit-learn__scikit-learn-10844: Pipeline estimators parameter validation fault.sphinx-doc__sphinx-8595: Python domain autodoc signature formatting error.
🔬 The Dual-Brain Cybernetic APR Architecture
Traditional LLM agents rely exclusively on multi-turn prompt loops, generating massive token overhead and hallucinating syntax formatting. Kronumos Kairos decouples cognition into two integrated computing cortices:
[Raw GitHub Issue Discussion]
│
▼
┌──────────────────────────────────┐
│ Sub-Cortex IssueDeNoiser │
│ - Excises human chatter/quotes │
│ - Tags user reproduction code │
│ - Extracts Core Signal Triad │
└────────────────┬─────────────────┘
│
[Cleaned Technical Specification]
│
▼
┌──────────────────────────────────┐
│ Procedural Cognitive Kernel │ ◄─── Procedural Seeds (<64 bytes)
│ - BoundaryCondition Invariants │ (Zero-DB L1 Cache Execution)
│ - DefensiveNullWrap / PopGuards │
└────────────────┬─────────────────┘
│
┌────────────────────┴────────────────────┐
▼ ▼
┌───────────────────────┐ ┌───────────────────────┐
│ CORTEX (Neural) │ │ SUBCORTEX (Deterministic)
│ Qwen2.5-Coder-7B │ ◄─────► │ Tree-sitter AST Slicer│
│ - 5-Step CoT Reason │ Dual-Key│ Auto-Bracket & Indent │
│ - Precise Code Hunk │ Consens.│ Merkle Causal Ledger │
└───────────────────────┘ └───────────────────────┘
- Issue De-Noiser: Strips human conversational chaff, extracting the core reproduction triad.
- Procedural Cognitive Kernel: Diagnoses invariants across 9 domains (Boundary, Defensive, Concurrency, etc.) without external database lookups.
- Dual-Key Consensus Gate: Requires simultaneous semantic approval and deterministic AST validation before admitting state changes.
- Auto-Bracket & Indentation Healer: Deterministically balances parentheses and enforces strict PEP 8 4-space block indentation.
⚡ Native Rust Sub-Cortex Runtime
This repository includes the precompiled native Linux x86_64 binary libtokenectomy_subcortex.so and Python C-ABI bridge tokenectomy_subcortex_rust.py.
from tokenectomy_subcortex_rust import RustSubCortex
subcortex = RustSubCortex()
# 1. Clean noisy issue descriptions (93.5% token reduction)
clean_spec = subcortex.denoise_issue(raw_github_issue)
# 2. Heal indentation drift and unbalanced brackets with sub-microsecond latency (5µs)
healed_code = subcortex.heal_indentation(candidate_code, base_indent=4)
💻 Quickstart: Running Inference
With Transformers:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "NadevA23/Kronumos-Kairos-v2"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
messages = [
{"role": "system", "content": "You are Kronumos Kairos, an expert autonomous program repair engine."},
{"role": "user", "content": "Fix the issue in the following function:\n\ndef safe_divide(a, b):\n return a / b"}
]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512, do_sample=False)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
Quantized GGUF (Llama.cpp / Ollama):
For quantized local execution on consumer hardware, visit NadevA23/Kronumos-Kairos-v2-GGUF.
📜 Citation
@article{daffa2026kronumos2,
author = {Muhammad Naufal Daffa},
title = {Kronumos 2 Kairos: Cost-Bounded Automated Program Repair via Dual-Brain Cybernetic Sub-Cortex on SWE-bench Verified},
journal = {Research Square},
year = {2026},
doi = {10.21203/rs.3.rs-11205335/v1},
url = {https://doi.org/10.21203/rs.3.rs-11205335/v1}
}
License: Apache 2.0
Maintained by: Tokenectomy Labs
- Downloads last month
- 1,317