Most Agentforce pilots that stall don't fail because of the AI. They fail because the org around the agent wasn't ready. The data couldn't answer the question, nobody owned the result, or permissions were too loose to put anything in front of a customer. Yesterday we covered what an agent is made of. Today is the practical follow-up: five signs your org is ready to run a pilot, and the red flags that mean you should fix something else first.
Why Readiness Beats a Good Idea
Every team can name something it would like an agent to handle. What separates a pilot that teaches you something from one that drags on for months is whether the conditions underneath are in place.
A pilot is an experiment. You want it to answer one question clearly: "Can an agent handle this job, in our org, well enough to be worth running in production?" If the data is unreliable, you won't know whether a bad answer came from the agent or the data. If nobody defined success, you can't tell whether the pilot worked. A readiness check exists to remove those confounding variables before you spend anything.
The five signs below aren't a maturity model. They're the conditions we look for before recommending a pilot. You don't need all five to be perfect, but you need all five to be true enough.
Your Data Can Already Answer the Question
An agent can only be as accurate as the records it grounds on. The test here is concrete. Take the use case you have in mind and ask a capable human who has never seen your org to answer ten typical requests using only what's in Salesforce.
If they can do it, your data is probably ready. If they keep saying "I'd need to check with someone" or "there are two accounts with this name", your agent will hit the same walls, only faster and in front of more people.
Things that indicate readiness:
- The fields the use case depends on are consistently populated, not optional fields that half the team skips
- There's one clear record for each customer, order, or case, not a tangle of duplicates
- Knowledge articles exist for the questions you expect, and someone keeps them current
- Status fields mean what they say ("Closed" actually means closed)
You don't need a perfectly clean org. You need the slice of data this one use case touches to be trustworthy. That's a much smaller job, and it's often a week of targeted cleanup rather than a program.
You Have One Narrow Repetitive Use Case
The best pilot candidates are boring. They happen many times a week, follow a recognizable pattern, and have a small number of correct outcomes. "Help customers with anything" isn't a use case. "Answer where-is-my-order questions for orders placed in the last 90 days, and hand off anything involving a refund" is.
Narrow scope is a feature, not a compromise. It lets you write a test set that covers the space, keeps the number of actions small, and gives you a clean before-and-after comparison.
A good way to check: can you write down twenty real examples of requests this agent would receive, along with the correct outcome for each? If you can, you have a use case. If the examples keep sprawling into new territory, you have a theme, and it needs narrowing before it becomes a pilot.
Repetition also matters for the business case. A task that happens twice a month won't justify the setup, testing, and monitoring an agent needs, however clever the agent is.
The Underlying Automation Already Works
As we covered on day one, agents mostly call Flows and Apex to get things done. So one of the strongest readiness signals is that the actions your use case needs already exist and already run reliably, even if today a person triggers them by hand.
If a working Flow already looks up order status or updates a delivery address, wrapping it as an agent action is a small step. If none of that exists, the "agent pilot" is really an automation project with an agent on the end. That's fine, but scope and budget it honestly.
Signs of automation maturity worth looking for:
- Core Flows have clear inputs and outputs and don't depend on hidden context from the screen a user happened to be on
- There's a sandbox where Flows are tested before deployment
- Someone can explain what each relevant Flow does without opening it
- Error handling exists, and failures don't just disappear silently
This connects to our own philosophy: if a problem repeats, it should become a system. Orgs that have already turned their repeated work into reliable automation are the ones where an agent can add a lot with little effort.
Permissions and Governance Are Under Control
An agent acts within a permission context. If your permission model is "most people are effectively admins", the agent inherits that sprawl, and you won't be able to say with confidence what it can and can't see.
Readiness here means you can answer three questions quickly:
- What should this agent be able to read? Specific objects and fields, not "the CRM".
- What should it be able to change? Ideally a short list, and always through defined actions.
- Who approves changes to its behavior? Topics, instructions, and actions will change over time, so someone has to sign those changes off.
You'll also want a view on sensitive data: which fields contain personal or regulated information, and whether the agent needs them at all. The Einstein Trust Layer helps with masking and auditing. But it works best when you've already decided what the agent should and shouldn't touch, rather than relying on it as your only line of defense.
Governance doesn't have to be heavy. For a pilot, one page listing the permission set, allowed actions, escalation path, and approver is usually enough, as long as it exists before the agent does.
Someone Owns It and Knows What Success Looks Like
This is the sign most often missing, and it's the one that decides whether a pilot turns into a decision.
An owner is a named person, usually from the business side, who cares about the outcome, can make scoping calls, and reviews results. Without one, pilots drift, and at the end nobody has the authority to say "ship it" or "stop".
A success metric is the number that owner will judge the pilot on. It should be measurable today, before the agent exists, so you have a baseline. Good examples:
- Share of in-scope requests resolved without a human
- Average handling time for a specific case type
- Time a rep spends preparing for a call
- Accuracy against a reviewed test set
We think this way about all our automation work. When we built a consolidated KPI reporting bot for founders who were logging into five tools every day, success was defined upfront and was easy to check: zero daily dashboard logins. A clear target like that makes it obvious whether a system is doing its job, and an Agentforce pilot deserves the same clarity.
Red Flags That Mean Not Yet
Some conditions mean a pilot will produce noise rather than answers. If you see these, fix them first.
| Red flag | What it causes in a pilot | What to do instead |
|---|---|---|
| Key fields are mostly empty or inconsistent | The agent gives wrong or vague answers, and you can't tell why | Targeted data cleanup for the use case's objects |
| The use case is "general assistant" | No test set, no baseline, endless scope creep | Pick one request type and narrow it |
| The actions it needs don't exist | The pilot turns into an unplanned automation build | Build and test the Flows first, then add the agent |
| Broad permissions nobody can explain | Real risk of exposing data; security blocks go-live | Create a dedicated, minimal permission set |
| No named business owner | Results go unreviewed, and no decision is made | Assign an owner before kickoff |
| Success is "see what it can do" | The pilot can't pass or fail | Agree on a baseline metric and a target |
| Leadership expects full automation on day one | The pilot is judged against the wrong bar | Frame it as a test with a human handoff built in |
None of these red flags mean "never". They mean the cheapest next step is to fix something outside the agent.
A Quick Self-Assessment
Score each sign from 0 to 2, where 0 means "not true", 1 means "partly", and 2 means "clearly true".
Data can answer the question 0 / 1 / 2
One narrow, repetitive use case 0 / 1 / 2
Underlying automation works 0 / 1 / 2
Permissions and governance in place 0 / 1 / 2
Named owner and success metric 0 / 1 / 2
These aren't scientific thresholds, just a practical way to read the result. If you're at 8 or above with no zeros, you're likely ready for a pilot. Between 5 and 7, you're close, and the low scores tell you exactly what to fix. If anything scores 0, especially ownership or data, deal with that before you configure a single topic.
The self-assessment gives you a direction. It won't catch the issues you don't know to look for: a Flow that quietly depends on a user's profile, a permission set that grants more than anyone realized, an integration that times out under load. That's what a structured readiness audit is for, and it's the subject of tomorrow's post.
Key Takeaways
- Most stalled pilots are caused by the org around the agent, not by the agent itself.
- The data only needs to be clean for the slice your use case touches, not for the whole org.
- A good pilot use case is narrow, repetitive, and easy to write twenty test examples for.
- If the Flows and actions don't exist yet, the pilot is really an automation project. Scope and budget it that way.
- A named owner and a baseline metric turn a pilot into a decision instead of a demo.
FAQs
Do we need a perfectly clean Salesforce org before piloting Agentforce?
No. You need the specific objects and fields your chosen use case depends on to be reliable. Focused cleanup of that slice is usually enough for a pilot.
Can IT own the pilot on its own?
IT should own the build, but the pilot needs a business owner who can make scope decisions and approve going to production.
What if we don't have any Flows for our use case yet?
Build and test those first. Debugging new automation and new agent behavior at once makes it hard to know what's broken.
How narrow is narrow enough?
Narrow enough that you can list around twenty realistic requests and the correct outcome for each, without the list spilling into unrelated areas.
Is an internal use case a safer first pilot than a customer-facing one?
Usually, yes. With call prep or case summaries, a person reviews the output before it reaches a customer, so you learn with lower risk.
Working on Something Similar
If you've scored yourself and aren't sure what the gaps mean, or you'd like a second opinion before committing budget, our Agentforce Opportunity Audit turns this checklist into a concrete recommendation for your org. See what's included on our Agentforce page, or reach out with the use case you're weighing up.


Comments
No comments yet. Questions and counterpoints are welcome.
Leave a comment