A working Agentforce proof of concept is a real milestone, and it's also where a lot of teams misjudge how much is left. The agent answers correctly in the demo, stakeholders are happy, and someone asks, "So we can switch it on next week?" Usually you can't, and the reason isn't that the POC failed. A POC and a production agent are built to answer different questions. This post lays out what actually changes between the two, dimension by dimension, so you can plan for that gap instead of discovering it.
Two Different Questions
A POC answers: "Can an agent do this job well enough, with our data and our actions, to be worth investing in?" It's an experiment. Its job is to reduce uncertainty as cheaply as possible, which means cutting corners on anything that doesn't affect the answer.
Production answers: "Can this agent do the job every day, for every user, safely, and keep doing it as things change?" That's an operational question. It brings in everything the POC was allowed to skip: edge cases, monitoring, rollback, ownership, and budget.
Neither is a smaller version of the other. A good POC deliberately leaves things out. The mistake is forgetting that you left them out.
The Comparison at a Glance
| Dimension | POC | Production |
|---|---|---|
| Scope | One narrow use case, a few topics, the core actions | The same use case, hardened, with every edge case handled or routed |
| Data | A representative sample, possibly a sandbox | Live data, at full volume, with a defined source of truth |
| Guardrails | Basic topic scope and a minimal permission set | Least-privilege access, verified identity, enforced limits, tested handoffs |
| Testing | A curated test set, scored manually | An automated regression suite run on every change, including adversarial cases |
| Monitoring | Someone reads the transcripts | Dashboards, alerts, conversation sampling, and error tracking |
| Rollout | Internal testers or a tiny group | Phased rollout with a kill switch and a rollback plan |
| Ownership | The project team | A named business owner plus a technical maintainer, with a change process |
| Cost | Fixed effort and low usage | Ongoing consumption, maintenance time, and a budget owner |
The rest of this post takes each row in turn.
Scope
In a POC you want the smallest scope that still answers the question. One request type, a couple of topics, and the three or four actions that matter. You might hand anything unusual straight to a person without trying to handle it.
In production, the scope of the use case often stays the same, but the scope of the behavior grows. Every way a real user can phrase, combine, or derail a request needs a defined response. Some get handled. Many get a clean, context-rich handoff. A few get a polite "that's not something I can help with". The work is in deciding which is which and making it consistent.
One trap to avoid: treating the jump to production as a chance to add new use cases. Harden the first one, ship it, and then expand. Doing both at once makes failures hard to attribute.
Data
POCs often run in a sandbox or against a sample of records, and that's fine for proving the concept. But production data is messier, bigger, and always changing. Records get merged, fields get renamed, and knowledge articles go stale.
Moving to production means confirming the agent grounds on the live source of truth for each fact. It means checking how fresh any synced or unified data is, and agreeing who keeps the relevant knowledge content current. If the POC relied on a few hand-cleaned records, assume the production data will expose new gaps. Budget time to find and fix them.
Guardrails
POC guardrails exist to keep the experiment safe: a restricted permission set, tight topic boundaries, and no actions that change anything important without review.
Production guardrails exist to keep the business safe at volume:
- Least-privilege access, reviewed by whoever owns security for the org, not inherited from a human role
- Identity verification as an action, so the agent can't share personal details until a check has passed
- Hard limits in actions, for example a Flow that refuses refunds above a threshold, rather than relying on instructions alone
- Tested handoffs, where escalations land in the right queue with the conversation context attached
- Trust Layer settings confirmed, including masking for sensitive fields and audit logging that someone actually reviews
The principle from day one applies here even more strongly. If it has to happen every time, enforce it in an action or a permission, not in an instruction.
Testing
A POC test set is usually a few dozen real requests, scored by people, mainly to show that the approach works. That's appropriate.
Production testing has a different job: catching regressions. Agent behavior is sensitive to small changes. Tweak one topic's instructions or one action description, and behavior can shift somewhere else. So production needs:
- a larger test set covering normal, edge, and adversarial requests
- a repeatable way to run it after every change, using Agent Builder's testing tools or your own harness
- clear pass criteria per request (right topic, right action, right outcome)
- a rule that changes don't deploy if the scores drop
Think of it the same way you'd think of a test suite for code. Nobody would ship an Apex change without tests. Agent changes deserve the same discipline.
Monitoring
In a POC, monitoring usually means a team member reading every transcript. At production volume that doesn't scale, and it misses slow drifts.
Production monitoring should tell you:
- how many conversations were resolved, handed off, or abandoned
- which actions are failing, and how often
- where users rephrase or repeat themselves, which is a sign the agent misunderstood
- whether usage, and therefore cost, is in line with expectations
Pair the metrics with regular sampled reviews of real conversations. Numbers tell you that something changed. Transcripts tell you why.
Rollout
A POC goes to internal testers. Production should go out in stages: a small percentage of traffic or one team first, then wider as the metrics hold.
We treat this the same way we'd treat any risky cutover. When we built a zero-downtime migration engine to move an application from MongoDB to normalized Postgres, the result was 0 seconds of application downtime and 95% faster query speeds. The approach that made it safe was batched streaming with the ability to verify each stage before moving on. Agent rollouts benefit from the same mindset: move in batches, verify each step, and always know how to roll back.
At a minimum, production rollout needs:
Stage 1 Internal users only -> verify handoffs and logging
Stage 2 Small slice of real traffic -> compare metrics to baseline
Stage 3 Wider rollout -> watch cost and failure rates
Always Kill switch + fallback path -> route back to humans instantly
Ownership
POCs are owned by the project team, who are motivated, close to the details, and temporary. Production needs owners who'll still be there in six months.
That usually means two roles. A business owner decides what the agent should do, reviews its performance, and approves scope changes. A technical maintainer handles topic, instruction, and action changes, runs the regression suite, and responds to failures. Add a lightweight change process, so edits to the agent's behavior get reviewed and tested rather than adjusted live in production because someone noticed an odd answer.
This is where many otherwise good agents slowly decay. Without clear ownership, small fixes pile up untested, and nobody notices the drift until users stop trusting the agent.
Cost
POC cost is mostly fixed: the time to build it, plus modest usage during testing. Production cost is ongoing, and it has three parts:
- Consumption, which scales with conversations and actions under Agentforce's consumption-based pricing
- Maintenance, meaning time spent updating knowledge, adjusting topics, and fixing actions as the business changes
- Monitoring and review, meaning the people time to read samples, triage issues, and report results
The POC should give you real numbers for steps per conversation and resolution rate. Use them to build a production cost model before rollout, and give that model a budget owner. Surprises here are what tend to stop agents after launch.
Planning the Gap Honestly
None of this means production is a mountain. For a well-scoped use case where the POC was done properly, hardening is a defined, bounded piece of work. The goal is simply to plan it as real work, with its own timeline, rather than as a formality after the demo.
A practical way to do that is to finish every POC with a short "production gap" list. Under each heading in the table above, write down what the POC deliberately skipped. That list becomes the scope of the production sprint. The next post in this series looks at the opposite problem: what happens when teams skip the POC entirely.
Key Takeaways
- A POC proves the agent can do the job. Production proves it can do it every day, safely, and keep doing so as things change.
- Guardrails, testing, monitoring, and rollout planning are where most of the production work lives.
- Treat agent changes like code changes: regression tests before every deployment.
- Roll out in stages, with a kill switch and a human fallback ready from day one.
- Production needs a business owner, a technical maintainer, and a cost model with a budget owner.
FAQs
Can we just promote our POC agent straight to production?
You can reuse its topics, actions, and test set, but you'll need to harden permissions, broaden testing, add monitoring, and plan a staged rollout. Treat the POC as the foundation, not the finished building.
How long does it take to go from POC to production?
It depends on how many gaps the POC deliberately left open and on the state of the underlying actions. Writing the production gap list at the end of the POC is the best way to get a realistic estimate.
What's the biggest difference between POC and production testing?
Purpose. POC testing shows the approach works. Production testing catches regressions, so it needs to be broader, automated, and run on every change.
Do we need new Flows for production?
Often you need to harden existing ones by improving error handling, enforcing limits, and returning clearer outputs. Sometimes you'll add new actions to handle cases the POC routed to people.
Who should own an agent in production?
A business owner accountable for its outcomes and a technical maintainer accountable for its behavior and changes. One person rarely covers both well.
Working on Something Similar
If you have a POC that works and a production go-live that feels further away than it should, our Agentforce Production Implementation Sprint is built to close that gap: hardening, testing, monitoring, and a staged rollout. Find out more on our Agentforce page, or contact us with where your agent is today.




Letters to the editor
No letters yet. Questions and counterpoints are welcome.