Introducing the Smaug Line: Abacus.AI's Open-Weight Models for Agentic AI
With three models built for different scales and workloads, the Smaug line represents our most significant step in making capable, production-ready agents accessible at scale.
At Abacus.AI, we have been focused on building self-improving agents that automate complex work processes. A persistent challenge that we kept running into was the cost and inefficiency of running large agentic loops with large contexts, repeated tool calls, and the compounding problem of prompt caching breaking down across time gaps.
To address this, we developed a fine-tuning methodology that combines human-curated, real-world agentic traces with synthetic data grounded in hard, challenging examples. Applied across a range of open-weight base models, the results are consistent: meaningful lifts in the benchmarks that actually matter for agents — agentic coding, real-world tool use, automation, and long-context reasoning and instruction following.
The Smaug line is a family of three open-weight models fine-tuned by Abacus.AI for agentic workloads, each built on a leading open-weight base at a distinct point on the capability–efficiency curve. The result is consistent, measurable performance gains across the benchmarks that matter for production agentic systems such as coding, tool use, automation, and long-context reasoning.
Use in RouteLLM API Hugging Face
Smaug Flash
Base model: DeepSeek V4 Flash 0731
Optimized for: Enterprise self-improving agents. Speed, cost, efficiency, and agentic performance at scale.
Optimized for: Enterprise self-improving agents. Speed, cost, efficiency, and agentic performance at scale.
Smaug Flash is the workhorse of the line, optimized for enterprise self-improving agents where speed, cost, efficiency, and agentic performance must coexist. Benchmarks show consistent gains across every major agentic workload. While DeepSeek Flash is a very efficient and robust model for frequently running agentic tasks, it is susceptible to spins and confusion with long context agentic tool use. Smaug Flash addresses this issue to make it work robustly and reliably while maintaining the cost and speed advantages of DeepSeek Flash.
Benchmarks
Benchmark Smaug Flash DeepSeek V4 Flash 0731 (base) Claude Sonnet 5
LiveBench overall 77.4 74.2 76.0
LiveBench agentic coding 61.1 46.8 59.4
DeepSWE 1.1 56.6 54.4¹ 55.0¹
AutomationBench (public 600, strict pass) 38.83 25.1¹ 34.67¹
NL2repo-bench 73.3 54.2¹ 66.34
LiveBench category profile
Smaug Flash LiveBench category profile against DeepSeek V4 Flash
Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.
Higher is better; highlighted = best in row (Sonnet 5 at xhigh effort). Evals were run by Abacus.AI under the same harness for every model unless marked: ¹ = reported score, not from our harness run.
Smaug Mini
Base model: Qwen3.8 27B
Optimized for: Multimodal use cases and smaller reasoning tasks
Optimized for: Multimodal use cases and smaller reasoning tasks
Smaug Mini targets multimodal use cases and smaller reasoning tasks, delivering strong agentic gains in a more compact, efficient package. Especially useful for one-off tasks that need multimodal capabilities. We improve upon Qwen3.8 27B to strengthen real-world agentic ability with this model.
Benchmarks
Benchmark Smaug Mini Qwen3.8-27B (base) Claude Sonnet 5 Claude Opus 4.6 GPT-5.6 Luna
LiveBench overall 76.9 75.3 76.0 74.5 73.6
LiveBench agentic coding 60.8 61.4 59.4 49.0 48.4
IFBench (prompt-level loose, 300) 82.0 79.5 69.3 63.0 75.33
AutomationBench (public 600, strict pass) 41.8 37.3 36.2 25.5 34.8
JobBench (official protocol, 65 tasks) 50.5 33.4 46.4 36.9¹ 43.3
NL2repo-bench 55.8 42.3 66.34 47.6¹ 52.9
LiveBench category profile
Smaug Mini LiveBench category profile against Qwen3.8 27B
Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.
Higher is better; highlighted = best in row. IFBench, AutomationBench and JobBench were run by Abacus.AI under the same harness for every model unless marked: ¹ = reported score, not from our harness run. LiveBench rows are the published livebench.ai 2026-06-25 leaderboard entries (Smaug Mini listed under the finetunes filter; Sonnet 5 at xhigh effort, Opus 4.6 at high effort).
Smaug Agentic
Finetuned on Kimi K3
Our largest model and clearest demonstration of how our fine-tuning technique scales
Our largest model and clearest demonstration of how our fine-tuning technique scales
Smaug Agentic is our largest model and clearest demonstration of how our fine-tuning technique scales. Applied to a frontier-scale multimodal base, Kimi K3, it holds its own and even leads in several key benchmarks.
Benchmarks
Benchmark Smaug Agentic Kimi K3 (base) GPT-5.6 Sol Claude Fable 5
GPQA Diamond 94.1 93.5 94.1 92.6
AA-LCR (long-context reasoning) 75.7 74.7 73.7 70.0
DeepSWE 69.9 67.5 73.0 70.0
Terminal-Bench 2.1 86.5 88.3 88.8 88.0
SciCode 60.8 58.7 56.1 60.2
LiveBench (Agentic Coding) 64.6 62.2 56.2 62.2
AutomationBench 31.0 30.8 29.7 29.1
MMMU-Pro 81.0 81.6 83.0 81.2
LiveBench category profile
Smaug Agentic LiveBench category profile against Kimi K3
Scores 0–100 per LiveBench category; overall is the mean of the seven category averages. Same harness, same questions (LiveBench 2026-06-25 release), paired against the base model.
Higher is better; highlighted = best in row. All models at max reasoning effort (Claude Fable 5 with fallback). Kimi K3, GPT-5.6 Sol and Claude Fable 5 reference scores are as published on the Kimi K3 model card; the Smaug Agentic LiveBench score is its published livebench.ai 2026-06-25 leaderboard entry.
Copyright © 2026 Abacus.AI. All Rights Reserved

Sign up to ChatLLM to proceed

Get more access to ChatLLM and unlock powerful AI Agent capabilities

Access to 100+ AI models including Fable 5.1, GPT 6 Astra and Seedream 2.0

Get Started
$10 $7
1st Month Discount First month then, $10/month
Models
100+ AI & Image Models
Vibe Code
Vibe Code Apps
General Purpose Agent
General Purpose Agent
CLI + CoWork
CLI + CoWork
SuperComputer
SuperComputer
Learn more