Paper notes
What Typhoon-S Teaches Us About Sovereign Agents
A bounded legal-agent result shows what small, locally grounded systems can do—and why the caveat is as important as the score.
Read the result precisely
In the NitiBench agentic setting reported by the Typhoon-S paper, the 4B Legal Agent scored 78.02%, versus 75.34% for GPT-5 + Agent. This is a bounded comparison inside a specific Thai legal reasoning environment. It is not evidence that a 4B model is generally more capable than GPT-5.
The precision matters because it reveals the real opportunity: local models can become highly effective when the domain, tools, reward, and evaluation are designed together.
Sovereignty is operational, not symbolic
A sovereign agent must understand local language and institutions, but it must also fit the deployment constraints of the organization using it. Open weights, reproducible post-training, local evaluation, and controllable infrastructure are all parts of the same system.
For sensitive workflows, the ability to run in a customer environment can matter as much as a benchmark score. It changes where data moves, how costs scale, and who can inspect the system.
The recipe is the product insight
The most useful takeaway from Typhoon-S is the recipe: start with a capable open model, preserve general abilities, add targeted training data, define task rewards, and evaluate inside the environment that matters.
Locatail applies that lesson to bounded retail conversations. The benchmark is not the destination. It is evidence for building Cartside around measured shopping work.
Source: Typhoon-S paper