Skip to content

Founder-led. Limited engagements.Talk to Us →

Agentforce Readiness Audit: What We Check and What Usually Breaks

The Agentforce readiness audit checklist we use: use cases, data, permissions, actions, testing, cost, and change management, plus what usually breaks.

By CodeITronics9 min readSalesforce · Agentforce · AI Audit · Governance
Agentforce Readiness Audit: What We Check and What Usually Breaks

A self-assessment tells you roughly where you stand. An audit tells you what will actually break. This post walks through the checklist we use when we assess an org for Agentforce: what we look at in each area, why it matters, and the patterns that tend to cause trouble later. None of these are rare edge cases. They're the ordinary ways a reasonable-looking plan runs into the reality of a live Salesforce org.

Why Audit Before You Build

Agentforce makes it fast to stand up something that looks like it works: a topic, a couple of actions, and a convincing conversation in Agent Builder within an afternoon. That speed is useful, but it hides problems. A demo on hand-picked inputs, run as an admin, in a sandbox with tidy sample data, tells you very little about how the same agent will behave with real users and real records.

An audit is the opposite of a demo. It looks for the reasons the agent won't work, while finding them on paper is still cheap.

The checklist is organized in the order we work through it, because each area depends on the ones before it.

Use-Case Selection

What we check. We list every candidate use case the team has in mind and score each one on volume, how varied the requests are, how many distinct outcomes there are, risk if the agent is wrong, and whether the actions it needs already exist. Then we look for the one or two that combine real value with a contained blast radius.

What usually breaks:

  • Starting with the most visible idea instead of the most suitable one. The use case leadership is excited about is often customer-facing, high-stakes, and broad. A better first move is usually internal and narrow, so the team learns how agents behave before the stakes rise.
  • Hidden branching. A use case that sounds simple, like "update a delivery address", turns out to have six policy exceptions that live in someone's head. If those exceptions aren't written down, the agent can't follow them.
  • Two use cases pretending to be one. "Handle billing questions" usually splits into explaining an invoice (informational) and changing a payment method (transactional). These carry very different risks and need different guardrails.

The output of this step is a short list, ideally with a clear first choice and the reasons the others were deferred.

Data and Grounding

What we check. For the chosen use case, we trace every piece of information the agent would need back to its source: which object, which field, which knowledge article, or which external system. Then we sample real records and check whether those fields are populated, consistent, and current. Where Data Cloud is in play, we check what's actually unified and how fresh it is, rather than relying on what the architecture diagram promises.

What usually breaks:

  • Answers that live in free-text fields. The information exists, but it's buried in a description or a notes field. Agents can read text, but grounding on inconsistent free text produces inconsistent answers.
  • Knowledge that's out of date. Articles that were accurate two product releases ago are worse than no article, because the agent will quote them confidently.
  • The same fact in two places. Order status in Salesforce and in the ERP, for example, with no agreement on which is right. The agent needs one source of truth per fact.
  • Sample data that hides the problem. Sandboxes often have cleaner, smaller data than production. We check against production-shaped data wherever we can.

Security Permissions and the Trust Layer

What we check. We look at the permission context the agent will run under, including which objects, fields, and records it can reach, and compare that with what the use case actually needs. We identify sensitive fields and confirm how the Einstein Trust Layer settings apply, such as data masking and audit logging. We also check how the agent verifies who it's talking to before it reveals anything specific to that person.

What usually breaks:

  • Testing as an admin. Everything works in Agent Builder because the person testing can see everything. Under a properly restricted agent user, actions start failing or returning empty results.
  • Permission sets copied from a human role. The agent ends up with the same access as a support rep, including fields and objects it has no reason to touch.
  • Identity checks handled by instructions. "Verify the customer before sharing order details" written as an instruction is a request, not a control. Verification should be an action with a clear pass or fail result that the topic depends on.
  • Assuming the Trust Layer covers everything. It handles important things like masking and auditing. It doesn't decide what the agent should be allowed to see in the first place. That's your permission design.

Actions and Integrations

What we check. We list every action the use case requires and classify each one: it exists and works, it exists but needs changes, or it needs to be built. For existing Flows and Apex, we check whether they can be called cleanly as agent actions. That means clear inputs and outputs, no dependence on a screen context, and sensible error handling. For external systems, we check authentication, latency, rate limits, and what happens when they're unavailable.

What usually breaks:

  • Flows built for screens. A Flow that assumes a user is on a particular record page, or that pops up a screen for confirmation, doesn't translate directly into an action an agent can call.
  • Silent failures. An action that fails without returning a useful error leaves the agent to improvise, and it may tell the user something happened when it didn't.
  • Slow external calls. An integration that takes several seconds is tolerable in a batch job and painful in a live conversation. Chaining a few of them makes it worse.
  • Vague action descriptions. If two actions sound similar, the agent will sometimes pick the wrong one. We rewrite descriptions so each action's purpose is unmistakable.

This is usually where the real effort hides. Configuring the agent is the smaller job; getting the actions solid is the bigger one.

Testing and Evaluation

What we check. We look at whether the team has, or can build, a realistic test set: actual requests drawn from cases, chats, or emails, each paired with the expected outcome. We check whether there's a repeatable way to run those tests after every change, and whether "correct" has been defined for each request type: the right action, the right answer, or the right handoff.

What usually breaks:

  • Testing only the happy path. The agent handles clean, polite, single-intent requests well. Real users send two questions in one message, leave out key details, change their minds, or try to push the agent off-topic.
  • No regression routine. A small change to one topic's instructions shifts behavior in another. Without re-running the full test set, nobody notices until users do.
  • Judging by impression. "It seemed pretty good" isn't an evaluation. We push for a scored test set, even a modest one, so changes can be compared objectively.

Cost and Consumption

What we check. Agentforce is priced on consumption, so we estimate usage from expected volume: conversations per day, typical steps per conversation, and which actions call the model. We then compare that with the value of the work being handled. We also look for places where deterministic work has been routed through the agent unnecessarily.

What usually breaks:

  • No usage estimate at all. The pilot runs, everyone likes it, and the first conversation about cost happens after the decision to scale. That's the wrong order.
  • Agents doing a Flow's job. Routine, structured work routed through the agent costs more and adds variability for no benefit.
  • Chatty designs. Agents that ask several clarifying questions where one lookup would do use more interactions and frustrate users.

We don't quote prices, since they depend on your Salesforce agreement. We give you a usage model to price with your account team.

Change Management

What we check. We identify who will work alongside the agent, such as support reps, sales reps, and operations staff, and how their day changes. We check how handoffs from agent to human will work in practice, who reviews agent conversations, and how feedback from frontline users gets back to whoever maintains the agent.

What usually breaks:

  • Handoffs without context. The agent escalates, and the human has to start the conversation again from scratch. That's worse than having no agent.
  • Teams who feel replaced rather than supported. Adoption suffers when people aren't told what the agent is for and what stays with them.
  • No feedback loop. Reps spot problems daily but have no simple way to report them, so the same issue repeats for weeks.

What an Audit Should Leave You With

A useful audit ends with decisions, not just observations. At minimum you should come away with:

1. A ranked use-case list, with a recommended first pilot
2. A data gap list for that pilot (field, issue, fix, owner)
3. A proposed permission set and escalation path
4. An action inventory (exists / modify / build)
5. A starter test set and a definition of "correct"
6. A usage model you can price with Salesforce
7. A rollout and feedback plan for the people affected

With that in hand, the next step is a scoped proof of concept that tests the riskiest assumptions first. That's where the rest of this series goes: tomorrow we look at what really changes between a POC and production.

Key Takeaways

  • An audit looks for the reasons an agent won't work, which a demo can't show you.
  • Most problems sit outside the agent itself: in data, permissions, existing Flows, and integrations.
  • Test under the agent's real permission context, never as an admin.
  • Controls that must hold every time belong in actions and permissions, not in instructions.
  • Estimate consumption before the pilot, not after everyone has decided to scale.

FAQs

How long does an Agentforce readiness audit take?

It depends on the org and how many use cases are on the table. A focused audit of a few candidates usually takes days to a couple of weeks, not months.

Do you need production access for an audit?

Read access to production metadata and representative data helps a lot, because sandboxes hide data issues. Otherwise we work from a full or partial copy sandbox and flag the limits.

What's the most common issue an audit uncovers?

Actions are the usual culprit. Existing Flows often weren't designed to be called by an agent, so they need cleaner inputs, outputs, and error handling before they're safe to use.

Can we run the audit ourselves?

Yes, this checklist is a good start. An outside audit helps most when the team is too close to the org to notice its own assumptions, or needs an independent recommendation.

Does an audit commit us to building an agent?

No. A good audit can conclude that a Flow, a data cleanup, or simply waiting is the better move. That's a valid and useful outcome.

Working on Something Similar

This checklist is essentially our Agentforce Opportunity Audit, applied to your org and your use cases, and it ends with a clear recommendation on what to pilot first. If that sounds useful, see the details on our Agentforce page or browse our other services, and get in touch when you're ready to talk it through.

Keep reading

5 Signs Your Org Is Ready for an Agentforce Pilot

Five practical signs your Salesforce org is ready for an Agentforce pilot, the red flags that mean not yet, and a quick self-assessment to score yourself.

What Salesforce Agentforce Actually Does (Beyond the Hype)

A plain explainer of Salesforce Agentforce: what an agent is made of, how it decides what to do, and when a simple Flow is the better choice.

Subscribe to the edition

One useful systems note a week from real Salesforce and automation builds. Unsubscribe in one click.

Letters

Letters to the editor

No letters yet. Questions and counterpoints are welcome.