Research
Introducing the Smaug Line: Abacus.AI's Open-Weight Models for Agentic AI
With three models built for different scales and workloads, the Smaug line represents our most significant step in making capable, production-ready agents accessible at scale.
At Abacus.AI, we have been focused on building self-improving agents that automate complex work processes. A persistent challenge that we kept running into was the cost and inefficiency of running large agentic loops with large contexts, repeated tool calls, and the compounding problem of prompt caching breaking down across time gaps.
To address this, we developed a fine-tuning methodology that combines human-curated, real-world agentic traces with synthetic data grounded in hard, challenging examples. Applied across a range of open-weight base models, the results are consistent: meaningful lifts in the benchmarks that actually matter for agents — agentic coding, real-world tool use, automation, and long-context reasoning and instruction following.
The Smaug line is a family of three open-weight models fine-tuned by Abacus.AI for agentic workloads, each built on a leading open-weight base at a distinct point on the capability–efficiency curve. The result is consistent, measurable performance gains across the benchmarks that matter for production agentic systems such as coding, tool use, automation, and long-context reasoning.
Smaug Flash
Base model: DeepSeek V4 Flash 0731
Benchmarks
| Benchmark | Smaug Flash | DeepSeek V4 Flash 0731 (base) | Claude Sonnet 5 |
|---|---|---|---|
| LiveBench overall | 77.4 | 74.2 | 76.0 |
| LiveBench agentic coding | 61.1 | 46.8 | 59.4 |
| DeepSWE 1.1 | 56.6 | 54.4¹ | 55.0¹ |
| AutomationBench (public 600, strict pass) | 38.83 | 25.1¹ | 34.67¹ |
| NL2repo-bench | 73.3 | 54.2¹ | 66.34 |
LiveBench category profile
| Category | Smaug Flash | Base | Delta | Score, 0–100 |
|---|---|---|---|---|
| Agentic Coding | 61.1 | 46.8 | +14.3 | |
| Instruction Following | 69.3 | 65.5 | +3.8 | |
| Mathematics | 90.1 | 86.8 | +3.3 | |
| Coding | 77.1 | 75.0 | +2.1 | |
| Language | 79.2 | 79.2 | +0.0 | |
| Data Analysis | 79.1 | 79.3 | -0.2 | |
| Reasoning | 86.2 | 86.6 | -0.4 | |
| Overall | 77.4 | 74.2 | +3.2 |
Smaug FlashDeepSeek V4 Flash 0731 (base)
Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.
Higher is better; highlighted = best in row (Sonnet 5 at xhigh effort). Evals were run by Abacus.AI under the same harness for every model unless marked: ¹ = reported score, not from our harness run.
Smaug Mini
Base model: Qwen3.8 27B
Optimized for: Multimodal use cases and smaller reasoning tasks
Benchmarks
| Benchmark | Smaug Mini | Qwen3.8-27B (base) | Claude Sonnet 5 | Claude Opus 4.6 | GPT-5.6 Luna |
|---|---|---|---|---|---|
| LiveBench overall | 76.9 | 75.3 | 76.0 | 74.5 | 73.6 |
| LiveBench agentic coding | 60.8 | 61.4 | 59.4 | 49.0 | 48.4 |
| IFBench (prompt-level loose, 300) | 82.0 | 79.5 | 69.3 | 63.0 | 75.33 |
| AutomationBench (public 600, strict pass) | 41.8 | 37.3 | 36.2 | 25.5 | 34.8 |
| JobBench (official protocol, 65 tasks) | 50.5 | 33.4 | 46.4 | 36.9¹ | 43.3 |
| NL2repo-bench | 55.8 | 42.3 | 66.34 | 47.6¹ | 52.9 |
LiveBench category profile
| Category | Smaug Mini | Base | Delta | Score, 0–100 |
|---|---|---|---|---|
| Mathematics | 89.7 | 86.2 | +3.5 | |
| Reasoning | 82.8 | 80.0 | +2.8 | |
| Language | 77.0 | 74.3 | +2.7 | |
| Data Analysis | 78.8 | 76.6 | +2.2 | |
| Instruction Following | 73.9 | 72.7 | +1.2 | |
| Coding | 75.4 | 75.7 | -0.3 | |
| Agentic Coding | 60.8 | 61.4 | -0.6 | |
| Overall | 76.9 | 75.3 | +1.6 |
Smaug MiniQwen3.8 27B (base)
Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.
Higher is better; highlighted = best in row. IFBench, AutomationBench and JobBench were run by Abacus.AI under the same harness for every model unless marked: ¹ = reported score, not from our harness run. LiveBench rows are the published livebench.ai 2026-06-25 leaderboard entries (Smaug Mini listed under the finetunes filter; Sonnet 5 at xhigh effort, Opus 4.6 at high effort).
Smaug Agentic
Finetuned on Kimi K3
Our largest model and clearest demonstration of how our fine-tuning technique scales
Benchmarks
| Benchmark | Smaug Agentic | Kimi K3 (base) | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|---|
| GPQA Diamond | 94.1 | 93.5 | 94.1 | 92.6 |
| AA-LCR (long-context reasoning) | 75.7 | 74.7 | 73.7 | 70.0 |
| DeepSWE | 69.9 | 67.5 | 73.0 | 70.0 |
| Terminal-Bench 2.1 | 86.5 | 88.3 | 88.8 | 88.0 |
| SciCode | 60.8 | 58.7 | 56.1 | 60.2 |
| LiveBench (Agentic Coding) | 64.6 | 62.2 | 56.2 | 62.2 |
| AutomationBench | 31.0 | 30.8 | 29.7 | 29.1 |
| MMMU-Pro | 81.0 | 81.6 | 83.0 | 81.2 |
LiveBench category profile
| Category | Smaug Agentic | Base | Delta | Score, 0–100 |
|---|---|---|---|---|
| Agentic Coding | 64.6 | 62.2 | +2.4 | |
| Data Analysis | 79.9 | 78.7 | +1.2 | |
| Coding | 82.5 | 81.4 | +1.1 | |
| Reasoning | 90.3 | 90.7 | -0.4 | |
| Instruction Following | 71.0 | 71.4 | -0.4 | |
| Mathematics | 83.9 | 84.4 | -0.5 | |
| Language | 84.4 | 85.5 | -1.1 | |
| Overall | 79.5 | 79.2 | +0.3 |
Smaug AgenticKimi K3 (base)
Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.
Higher is better; highlighted = best in row. All models at max reasoning effort (Claude Fable 5 with fallback). Kimi K3, GPT-5.6 Sol and Claude Fable 5 reference scores are as published on the Kimi K3 model card; the Smaug Agentic LiveBench score is its published livebench.ai 2026-06-25 leaderboard entry.