Skip to content
← All posts

7 September 2026

4 min read

Written by

Clément Lacaille

Clément Lacaille

Founder, Tech-Bharat

About the author
Studio notes

Studio notes — making AI agents reliable in production

This week the agent fleet shipped no big new feature, but three architectural choices spotted across ongoing engagements are worth sharing: caching AI email triage instead of recomputing it on every click, keeping a Qualiopi compliance agent “on a leash” through traceable tool calls, and forcing an insurance coverage check into a strict JSON schema instead of a prose summary. Three transferable lessons for any owner or CIO putting AI agents into production.

AI email triage: a result you cache, not an answer you re-ask on every click

This week the agent fleet delivered no headline feature at any client, but three design choices noticed while reviewing ongoing engagements are worth sharing — they apply to any SME weighing AI email automation, or AI agents in production more broadly. First engagement: at a courier company, the agent that sorts the shared inbox (urgent client issue, driver problem, invoice, internal note…) does not re-query a language model every time a message is opened. Each email is scored once, against a strict output schema, the result is written to a database, and every later read hits that database — not the model. If the model or the database is briefly unreachable, the agent degrades gracefully instead of blocking the inbox. Treating a model’s output as an expensive side-effect you cache, rather than a live dependency you query on demand, is what separates a robust AI email automation from a demo that breaks on the first restart. The choice matters more given how much task surface a language model can realistically take on: a 2023 study by Eloundou, Manning, Mishkin and Rock, GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models, estimates that around 80% of the US workforce could have at least 10% of their tasks affected by this kind of model — a figure that explains why sorting a shared inbox is worth automating at all, and makes doing it reliably, rather than re-asking a model on every click, all the more necessary.

A compliance agent that can only move forward by calling tools

Second engagement, at a vocational training organization preparing for its quality-certification audit: the chat assistant that walks an internal auditor through the audit, indicator by indicator, is deliberately kept “on a leash.” It can only progress the audit by calling explicit tools — save an answer, attach evidence, advance to the next indicator — that write to a relational database acting as the sole source of truth, never by improvising freely in the conversation. For a Qualiopi compliance workflow, that trades a bit of conversational fluidity for a fully traceable, replayable trail — which matters more than eloquence when an external auditor asks to see exactly how an answer was reached. It is precisely the principle validated by ReAct, an agent architecture proposed by Yao, Zhao, Yu, Du, Shafran, Narasimhan and Cao in ReAct: Synergizing Reasoning and Acting in Language Models: by interleaving explicit reasoning with calls to external tools, rather than letting the model produce free-form text, the authors show the resulting trajectories are both more reliable — fewer hallucinations and less error propagation — and easier for a human to verify after the fact.

Replacing a prose summary with a strict JSON schema, for an agent that actually runs in production

Third engagement, at an insurance brokerage that compares a policy against the underlying contract to surface coverage gaps: rather than asking the agent for a prose summary, its response is forced through a strict JSON schema — fixed categories, severity levels, no free-text fields — directly usable as a table rather than text to be re-parsed. That choice is not free: a 2024 study by Tam, Wu, Tsai, Lin, Lee and Chen, Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models, finds that overly constrained output formats can hurt a model’s reasoning ability — but that the same constraint improves accuracy on classification tasks, exactly the shape of a coverage check between two contracts. The lesson holds for any owner or CIO putting an AI agent into production: the output format is not an implementation detail, it is a trade-off between reasoning richness and the reliability of what comes out of the black box — and the right choice depends on the task, not on an aesthetic preference for natural language.

Free resource

The self-assessment grid: 20 tasks AI can automate

Sales, admin, support, operations: the 20 tasks AI agents already handle in SMEs — with, for each one, the tell-tale sign that your team is concerned.

Read next