Skip to content

Research

Introducing the Smaug Line: Abacus.AI's Open-Weight Models for Agentic AI

With three models built for different scales and workloads, the Smaug line represents our most significant step in making capable, production-ready agents accessible at scale.

At Abacus.AI, we have been focused on building self-improving agents that automate complex work processes. A persistent challenge that we kept running into was the cost and inefficiency of running large agentic loops with large contexts, repeated tool calls, and the compounding problem of prompt caching breaking down across time gaps.

To address this, we developed a fine-tuning methodology that combines human-curated, real-world agentic traces with synthetic data grounded in hard, challenging examples. Applied across a range of open-weight base models, the results are consistent: meaningful lifts in the benchmarks that actually matter for agents — agentic coding, real-world tool use, automation, and long-context reasoning and instruction following.

The Smaug line is a family of three open-weight models fine-tuned by Abacus.AI for agentic workloads, each built on a leading open-weight base at a distinct point on the capability–efficiency curve. The result is consistent, measurable performance gains across the benchmarks that matter for production agentic systems such as coding, tool use, automation, and long-context reasoning.

Smaug Flash

Base model: DeepSeek V4 Flash 0731

Optimized for: Enterprise self-improving agents. Speed, cost, efficiency, and agentic performance at scale.

Smaug Flash is the workhorse of the line, optimized for enterprise self-improving agents where speed, cost, efficiency, and agentic performance must coexist. Benchmarks show consistent gains across every major agentic workload. While DeepSeek Flash is a very efficient and robust model for frequently running agentic tasks, it is susceptible to spins and confusion with long context agentic tool use. Smaug Flash addresses this issue to make it work robustly and reliably while maintaining the cost and speed advantages of DeepSeek Flash.

Benchmarks

Benchmark Smaug Flash DeepSeek V4 Flash 0731 (base) Claude Sonnet 5
LiveBench overall 77.4 74.2 76.0
LiveBench agentic coding 61.1 46.8 59.4
DeepSWE 1.1 56.6 54.4¹ 55.0¹
AutomationBench (public 600, strict pass) 38.83 25.1¹ 34.67¹
NL2repo-bench 73.3 54.2¹ 66.34

LiveBench category profile

Category Smaug Flash Base Delta Score, 0–100
Agentic Coding 61.1 46.8 +14.3
Instruction Following 69.3 65.5 +3.8
Mathematics 90.1 86.8 +3.3
Coding 77.1 75.0 +2.1
Language 79.2 79.2 +0.0
Data Analysis 79.1 79.3 -0.2
Reasoning 86.2 86.6 -0.4
Overall 77.4 74.2 +3.2

Smaug FlashDeepSeek V4 Flash 0731 (base)

Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.

Higher is better; highlighted = best in row (Sonnet 5 at xhigh effort). Evals were run by Abacus.AI under the same harness for every model unless marked: ¹ = reported score, not from our harness run.

Smaug Mini

Base model: Qwen3.8 27B

Optimized for: Multimodal use cases and smaller reasoning tasks

Smaug Mini targets multimodal use cases and smaller reasoning tasks, delivering strong agentic gains in a more compact, efficient package. Especially useful for one-off tasks that need multimodal capabilities. We improve upon Qwen3.8 27B to strengthen real-world agentic ability with this model.

Benchmarks

Benchmark Smaug Mini Qwen3.8-27B (base) Claude Sonnet 5 Claude Opus 4.6 GPT-5.6 Luna
LiveBench overall 76.9 75.3 76.0 74.5 73.6
LiveBench agentic coding 60.8 61.4 59.4 49.0 48.4
IFBench (prompt-level loose, 300) 82.0 79.5 69.3 63.0 75.33
AutomationBench (public 600, strict pass) 41.8 37.3 36.2 25.5 34.8
JobBench (official protocol, 65 tasks) 50.5 33.4 46.4 36.9¹ 43.3
NL2repo-bench 55.8 42.3 66.34 47.6¹ 52.9

LiveBench category profile

Category Smaug Mini Base Delta Score, 0–100
Mathematics 89.7 86.2 +3.5
Reasoning 82.8 80.0 +2.8
Language 77.0 74.3 +2.7
Data Analysis 78.8 76.6 +2.2
Instruction Following 73.9 72.7 +1.2
Coding 75.4 75.7 -0.3
Agentic Coding 60.8 61.4 -0.6
Overall 76.9 75.3 +1.6

Smaug MiniQwen3.8 27B (base)

Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.

Higher is better; highlighted = best in row. IFBench, AutomationBench and JobBench were run by Abacus.AI under the same harness for every model unless marked: ¹ = reported score, not from our harness run. LiveBench rows are the published livebench.ai 2026-06-25 leaderboard entries (Smaug Mini listed under the finetunes filter; Sonnet 5 at xhigh effort, Opus 4.6 at high effort).

Smaug Agentic

Finetuned on Kimi K3

Our largest model and clearest demonstration of how our fine-tuning technique scales

Smaug Agentic is our largest model and clearest demonstration of how our fine-tuning technique scales. Applied to a frontier-scale multimodal base, Kimi K3, it holds its own and even leads in several key benchmarks.

Benchmarks

Benchmark Smaug Agentic Kimi K3 (base) GPT-5.6 Sol Claude Fable 5
GPQA Diamond 94.1 93.5 94.1 92.6
AA-LCR (long-context reasoning) 75.7 74.7 73.7 70.0
DeepSWE 69.9 67.5 73.0 70.0
Terminal-Bench 2.1 86.5 88.3 88.8 88.0
SciCode 60.8 58.7 56.1 60.2
LiveBench (Agentic Coding) 64.6 62.2 56.2 62.2
AutomationBench 31.0 30.8 29.7 29.1
MMMU-Pro 81.0 81.6 83.0 81.2

LiveBench category profile

Category Smaug Agentic Base Delta Score, 0–100
Agentic Coding 64.6 62.2 +2.4
Data Analysis 79.9 78.7 +1.2
Coding 82.5 81.4 +1.1
Reasoning 90.3 90.7 -0.4
Instruction Following 71.0 71.4 -0.4
Mathematics 83.9 84.4 -0.5
Language 84.4 85.5 -1.1
Overall 79.5 79.2 +0.3

Smaug AgenticKimi K3 (base)

Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.

Higher is better; highlighted = best in row. All models at max reasoning effort (Claude Fable 5 with fallback). Kimi K3, GPT-5.6 Sol and Claude Fable 5 reference scores are as published on the Kimi K3 model card; the Smaug Agentic LiveBench score is its published livebench.ai 2026-06-25 leaderboard entry.