Most AI agent projects don't fail in production. They fail earlier, quietly, in the weeks between a promising demo and a launch date that keeps moving. The demo works on five hand-picked examples, everyone gets excited, and then the team discovers it has no idea whether the agent is actually good, what it costs, or who catches it when it's wrong.
The patterns behind those stalls are remarkably consistent. None of them are about the model being too weak. They're about scope, data, evaluation, and operations, which are the same things that sink any software project, only harder to see because an LLM makes a rough prototype look finished.
The demo trap
An agent prototype is cheap to build. Wire a model to two or three tools, write a system prompt, try a few inputs, and it looks like magic. That's the trap. The gap between "it worked when I tried it" and "it works on what real users send" is much larger for agents than for traditional software, because the behavior isn't specified anywhere. It emerges from the prompt, the model, the tools, and the data together.
So the demo creates false confidence, and the unglamorous work of defining success, measuring it, and building the operational shell around the model gets pushed to "after launch." Then launch never comes, because nobody can honestly say it's ready.
Here are the seven patterns we see most often, and what to do instead.
Vague scope
"Build an agent that handles customer emails" is not a scope. It's a wish. Which emails? All of them, including legal threats and refund disputes? What does "handle" mean: draft a reply, send it, update the CRM, close the ticket?
Vague scope causes two problems. The prompt becomes a sprawling list of edge cases, so the agent does nothing well. And nobody can agree when it's done.
What works instead is narrowing to one job with a clear input, a clear output, and a clear boundary:
| Vague | Scoped |
|---|---|
| Handle customer emails | Classify inbound support emails into 8 intents and draft replies for the 3 most common |
| Help sales with research | For each new inbound lead, produce a one-page company brief and a suggested first line |
| Automate reporting | Post a daily summary of 6 named metrics to one Slack channel at 9am |
Notice the scoped versions are almost boring. That's the point. Boring scope ships. You can widen it later once the first job is reliable.
No evaluation set
In our experience, this is the clearest dividing line between agent projects that ship and ones that don't. If you can't measure whether a change made the agent better or worse, you can't improve it, and you can't confidently say it's ready.
It doesn't need to be fancy. Start with 50 to 200 real inputs from historical data, each paired with what a good output looks like. For a classifier that's the correct label. For a drafting agent it might be a rubric: did it cite the right policy, did it avoid promising anything it can't deliver, is the tone right.
Run every prompt, model, and tool change against that set before shipping. Without it, teams fall into prompt whack-a-mole: fix one reported bad output, silently break three cases nobody re-tested.
A few practical notes:
- Include hard cases on purpose: ambiguous inputs, missing fields, angry customers.
- Include cases where the right answer is "I don't know" or "escalate." An agent that never abstains is an agent you can't trust.
- Version the set alongside the prompts.
Bad or ungrounded data
An agent is only as good as what it can see. If it answers support questions from a knowledge base that's two years stale, it will confidently give two-year-old answers. If it researches leads from a CRM full of duplicates and blank fields, its research will be garbage with good grammar.
The related failure is ungrounded generation: asking the model to answer from general knowledge when the answer should come from your data. That's where hallucinations live.
The fix has two parts. Ground the agent in retrieval: fetch the relevant documents or records, put them in context, and instruct the model to answer only from them. And treat data quality as part of the project, not someone else's prerequisite. Often the first real deliverable is a cleaned-up knowledge base, and that pays off whether or not the agent ever ships.
Using an agent where a workflow would do
Not everything needs an agent. An agent, in the sense of a model that decides which steps to take and in what order, is the right tool when the path genuinely varies from input to input. When the path is known, a workflow is cheaper, faster, more predictable, and far easier to debug.
A useful test: can you draw the process as a flowchart with a handful of branches? If yes, build the flowchart. Use the LLM at the specific steps that need language understanding, like classifying an email or extracting fields from a PDF, and let deterministic code handle the rest.
Agent-shaped: input -> model decides -> tool? -> model decides -> tool? -> ... -> output
Workflow-shaped: input -> classify (LLM) -> route (code) -> extract (LLM) -> write to CRM (code)
Most systems we build at CodeITronics are the second shape. Starting with a fully autonomous agent for a workflow-shaped problem means fighting non-determinism you never needed.
No observability
When an agent produces a bad output in production, the first question is "why?" If you can't answer it, you can't fix it.
Observability means logging every run in enough detail to reconstruct it: input, retrieved context, each model call and response, each tool call and result, final output, latency, and tokens. Tracing tools exist, but even a well-structured table in your own database beats nothing.
Teams skip this because it isn't visible in a demo. Then someone asks why the agent told a customer the wrong thing, and the honest answer is "we don't know." That moment tends to end pilots.
Two things to log that people forget: the version of the prompt and the version of the retrieved data. A bad answer is often caused by a document that was updated yesterday, not by the model.
No human-in-the-loop design
"We'll have a human check it" is not a design. Which outputs get checked? By whom? In what tool? What happens to the outputs they reject, and does anyone learn from them?
Good human-in-the-loop design decides up front where the agent acts alone and where it needs approval. Usually that's driven by two things: the model's confidence and the cost of being wrong.
| Confidence | Low-cost action | High-cost action |
|---|---|---|
| High | Act automatically | Act, but log and sample for review |
| Low | Act with a flag, or ask | Route to a human with the draft attached |
The review experience matters as much as the rules. If reviewing a draft takes longer than writing it, people stop reviewing. Put the draft, its source context, and approve/edit/reject in the tool the reviewer already uses, and feed rejections back into the evaluation set.
Cost surprises
A prototype that costs a few cents per run looks free. Then it meets production volume, or someone adds a retry loop, or an agent gets stuck calling the same tool repeatedly, and the bill arrives.
Agent costs depend on behavior. A model that takes three steps on one input might take twelve on another, and long contexts and verbose tool outputs multiply tokens.
What helps:
- Estimate cost per task early, using your evaluation set, not a single example. Look at the spread, not just the average.
- Set hard limits: maximum steps per run, maximum tokens per run, timeouts on tool calls.
- Use smaller, cheaper models for simple steps like classification and reserve larger models for the steps that need them.
- Cache what you can. Many inputs repeat, and many retrieved documents don't change between runs.
- Put cost on a dashboard next to quality. If you only watch one, you'll be surprised by the other.
A pre-launch checklist
Before calling an agent ready, we want a clear yes to each of these:
- Can we describe the job in one sentence, with a clear input, output, and boundary?
- Do we have an evaluation set built from real data, including hard cases and abstain cases?
- Is the agent grounded in current, reasonably clean data, and do we know who keeps it current?
- Have we justified each autonomous decision, or could parts of it be a fixed workflow?
- Can we reconstruct any single run from logs?
- Do we know which outputs a human reviews, and is reviewing faster than doing?
- Do we know the cost per task, and are there hard limits on runaway runs?
None of these require a breakthrough. They require treating the agent as a production system from day one.
Key Takeaways
- Agent projects usually stall because of scope, data, evaluation, and operations, not model capability.
- An evaluation set built from real inputs is the single most important asset in an agent project.
- If the process fits a simple flowchart, build a workflow with LLM calls at the ambiguous steps instead of a fully autonomous agent.
- Design human review deliberately, based on confidence and cost of error, and make reviewing faster than doing.
- Measure cost per task on realistic inputs early and enforce hard limits on steps and tokens.
FAQs
How big does an evaluation set need to be?
For a narrow task, 50 to 200 well-chosen real examples is a practical start. Grow it every time production reveals a case you hadn't covered.
When is a fully autonomous agent the right choice?
When the steps genuinely vary by input and can't be enumerated in advance, like open-ended research. Even then, constrain its tools and cap its steps.
Do we need a vector database to ground an agent?
Not always. For structured data, a normal database query is better. Vector search helps when you need relevant passages from unstructured text like help articles or past tickets.
How do we stop the agent from making things up?
Retrieve the source material, instruct the model to answer only from it, require it to say when the answer isn't there, and test that behavior in your evaluation set.
What's the fastest way to get an agent project unstuck?
Narrow the scope to one job, build an evaluation set for it, and measure where you actually are. Stuck projects usually don't know how far from ready they are, and that uncertainty is the real blocker.
Working on Something Similar
If you have an agent prototype that demos well but hasn't made it to production, the gap is usually one or two of the patterns above. CodeITronics builds AI agents and workflow automation as production systems, with evaluation, observability, and human review designed in from the start. You can see how we approach it on our services page, browse past builds on our work page, or get in touch to talk through where your project is stuck.




Letters to the editor
No letters yet. Questions and counterpoints are welcome.