Ilustración editorial para Cómo elegir una solución de IA para atención al cliente según la tarea, los datos y el riesgo
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

Start with the outcome you want to achieve

Choosing a customer service AI solution does not start with selecting a model. It starts with describing the work you want to change and the outcome that will count as satisfactory. “Improve support” is too broad to guide a decision: it could mean classifying messages with less intervention, finding up-to-date information more easily, preparing drafts for an agent, or completing an action requested by a customer. Each goal calls for different data, controls, and tests.

Define the use case by answering five questions: What requests come in, through which channels and in which languages? Who receives each request today? What outcome should the system produce? Which exceptions are out of scope? And who takes over when the system cannot resolve a case? Also define what it means to resolve a contact: sending a response is not necessarily the same as providing a correct solution, and closing a ticket does not, by itself, show that the customer's need has been met.

Separate the operational goal from the automation goal. Reducing the time it takes to prepare a response can be useful even if a person still reviews it. By contrast, reducing the number of agents involved should not count as success if incorrect responses, corrections, or repeat contacts increase. These are recommendations for designing an evaluation; they should be adapted to each organization's policies, service, and commitments.

Initial use-case brief

Complete these points before comparing technologies:

  1. 01Describe one specific request type and how it is identified.
  2. 02Record the channel, languages, and system in which the request is managed.
  3. 03Define the correct outcome and the errors that would be unacceptable.
  4. 04State which cases are excluded and how exceptions are routed to a person.
  5. 05Identify who can review changes to rules, sources, and responses.
02

Classify the task before choosing the technology

Several tasks may arise in the same conversation, but it is useful to evaluate them separately. Classification and routing assign a category or destination; information retrieval finds relevant content; drafting proposes a response; summarization condenses an exchange; and taking an action changes something outside the conversation, such as an account or a ticketing system. The closer a task gets to changing an account's state, the more important authorization, verification, and the ability to reverse the change become.

Not all of these tasks need text generation. If a request can be identified through stable fields and clear rules, conventional automation may be easier to test and maintain. A bounded classifier may be useful when language varies in ways that rules do not handle well; a generative system may be appropriate if a question needs to be interpreted and a response formulated. The choice depends on context: the fact that an interaction takes place in chat does not, by itself, mean a generative model is the best option.

The available sources discuss AI in customer service as an area for automation and assistance, but the information provided does not independently compare the effectiveness of each pattern or support a claim that one is better than another. Treat the comparison as a hypothesis to test against representative requests from your own service, not as a promised result.

Task map and initial pattern

The suggested pattern is a starting point for evaluation, not a guarantee of effectiveness.

TaskDesired outcomeMinimum pattern worth evaluatingPrimary control
Classify and routeAssign a category, queue, or priorityRules, bounded classification, or a combinationMeasure errors by category and allow assignments to be corrected
Find informationLocate relevant, up-to-date contentSearch an authorized document collectionCheck that the response is supported by applicable content
Draft or summarizePrepare text for an agentReviewable draft or conversation summaryHuman review and checks for omissions or unsupported claims
Take an actionChange a field or complete an operationA limited tool with authorization and confirmationVerify the action, its scope, and the response to failures
03

Match delegation to the consequences of an error

The level of autonomy should not depend only on technical difficulty. Ask what could happen if a response is wrong or incomplete, or is applied to the wrong account. General information that is easy to check and has no lasting effects presents a different kind of risk from a decision that changes an account, affects a payment, or determines access to a service. Reversibility, the time available to correct a problem, and a person's actual ability to intervene also matter.

As an operational framework, distinguish three levels. At the first, the system organizes or searches for information but does not respond to the customer or change their account. At the second, it prepares a response or recommendation that a person validates before sending. At the third, it can respond or take actions within explicit limits. Moving from one level to the next requires evidence that the previous level meets agreed criteria, as well as suitable controls for the next one; greater automation is not necessary to demonstrate value.

Keep a visible route for abstaining and escalating. If information is missing, a request is ambiguous, sources conflict, or a customer raises a dispute, the system should be able to stop short of proposing an automated resolution. In sensitive cases, human review must mean more than keeping a record: the reviewer needs enough context, authority to correct the result, and a process for stopping the action.

The Spanish Data Protection Agency includes a guide on agentic AI in its materials, a relevant area when a system can perform tasks. The available note confirms that general relevance but does not provide enough detail to attribute specific authorization, review, or design rules to the guide. The controls in this section are therefore proposed decision criteria and should be checked against the policies and obligations that apply to the use case.

04

Examine the data before connecting it

Inventory the sources each pattern would use: incoming messages, conversation histories, account records, help documentation, internal policies, or product information. For each source, note who is responsible for it, when it was updated, who is permitted to access it, and whether it contains personal data, confidential information, or information that should not leave the intended environment. Do not assume a document collection is safe or up to date simply because support already uses it.

Minimize the material needed for the task. Classifying an inquiry may not require transferring an entire account history; answering from documentation may only require retrieving relevant content rather than including large amounts of data in a prompt. Consider whether identifying information can be excluded, hidden, or replaced, and check current primary documentation to understand how any external service under consideration handles data. The information available here does not support claims about specific providers' retention, security, or data-use policies.

Also document usage conditions and internal responsibilities. The PwC good-practice guide identified among the sources touches on privacy and choices about data collection and use, but the evidence available is a fragment, not a full review of its recommendations. Treat it as a signal that privacy should be part of the evaluation, not as a substitute for legal, security, or policy review within the organization.

If the knowledge changes frequently, assign a person or team to review documents, effective dates, and withdrawn content. Design a test that checks not only whether the system finds an answer, but also whether it distinguishes current, incomplete, and conflicting information. If you cannot establish which source is authoritative for each topic, it is too early to ask the system to respond without review.

Data and source checks

  1. 01List every data source and the specific purpose for which it would be used.
  2. 02Confirm permissions, ownership, update date, and usage restrictions.
  3. 03Identify sensitive information and decide whether it can be excluded or reduced.
  4. 04Check how data would be handled by each external service without assuming conditions that are not documented.
  5. 05Define how obsolete content will be removed and who approves changes to the knowledge base.
05

Choose the smallest sufficient pattern for the task

For classification and routing, compare existing rules with a classification-based alternative. Evaluate both on varied messages, including those with multiple issues, typos, or insufficient information. The question is not only what proportion is assigned correctly, but also which categories account for the errors, how those errors are corrected, and what happens when the system is not confident enough. If the rules already solve the problem clearly and at an acceptable operating cost, replacing them with text generation may add complexity without demonstrated value.

For finding answers in documentation, first define which collection can be searched and how you will recognize a supported answer. A useful test includes questions answered in the documentation, questions outside its scope, and questions where the content is outdated or contradictory. Check that the system can identify the source it used and abstain when there is not enough evidence. Showing a citation or excerpt does not, by itself, guarantee that the answer is correct: a person must verify whether the source fits the question and is still current.

For drafting or summarizing, treat the result as a draft. Define what an agent may change before sending it, and review for context errors, omissions, inappropriate tone, and unauthorized promises. If the team cannot spend time reviewing and correcting the drafts, the assumed savings from producing more of them may disappear. Measure usefulness through work actually saved and quality, not by counting generated text.

For taking actions, limit the available operations and separate interpreting the request from authorizing a change. Before an action, check identity and permissions according to the organization's current procedures. For an initial test, consider reversible actions with a limited scope, and log what was requested, authorized, and carried out. Post-action verification should confirm the actual state of the system, rather than relying only on the message returned by the model.

These patterns can be combined, but doing so multiplies the points that need testing: information retrieval, interpretation, generation, integration, and execution. Start with a small workflow, record where it fails, and add components only when they address an observed gap.

06

Compare cost and operating conditions

Compare cost per resolved case, not just the price of a call or license. Where applicable, include integration with the ticketing system, infrastructure, maintenance of rules or documentation, monitoring, human review time, error correction, and handling of escalated cases. Measure cost using a volume and definition of resolution comparable to the current process; otherwise, the comparison may be between different tasks.

Day-to-day operations also affect the choice. Check whether the integration can show an agent the context they need, whether latency fits the channel, who maintains the knowledge base, what happens if a dependency becomes unavailable, and who is responsible for reviewing results. If there is after-hours coverage, define clearly what can be resolved without a person and what must wait or be escalated. Do not assume that “available” means “resolved.”

The business guide from Softeng identified among the sources covers use cases, data governance, security, and return. That description supports including those dimensions in the evaluation, but it does not, by itself, provide comparable cost or performance data for a particular team. Similarly, IBM's general material on AI in customer service can provide contextual guidance on the topic, but not independent proof of effectiveness or expected results.

Initial decision table

Use this table to organize questions and choose a test scope; it does not replace a technical, privacy, or applicable-obligations assessment.

ProblemData and evidenceImpact of an errorBudget and operationsCautious starting point
Inquiry routingLabeled messages and agreed categoriesDelay or assignment to the wrong queueReview ticketing integration and label correctionCompare rules with bounded classification
Documentation-based answersAuthorized, current, relevant contentIncorrect information or information applied to a different caseMaintain sources, measure escalations, and plan for reviewSearch and draft with visible supporting evidence
Preparing responsesNecessary context and representative examplesOmission, unsupported claim, or inappropriate toneCount review and correction timeDrafts not sent automatically
Action on an accountRequest, permissions, and verifiable stateIncorrect or difficult-to-reverse changeIntegration, logging, confirmation, and recoverySimulation or reversible action with approval
07

Design a test that can change the decision

Before deployment, prepare a sample that reflects real requests but also includes difficult cases: ambiguous messages, out-of-scope requests, policy changes, conversations with multiple needs, and missing key information. Define who reviews each output and what criteria they will use. If you test only easy examples selected to show the system at its best, the results will not tell you much about everyday use.

Agree in advance on the metrics and thresholds that will lead you to proceed, review, or stop the test. Depending on the task, relevant measures might include correct classification by category, responses supported by current sources, appropriate escalations, serious errors, time to a useful response, review time, and cost per resolved case. Do not turn a single general metric into an automatic decision: a favorable average can hide failures concentrated in one category or group of requests.

Compare the test with the current process using the same types of requests and a shared definition of a correct outcome. Record errors and their consequences, not just how often they occur. Distinguish a failure that an agent can correct before sending from an action that has already changed an account. When incidents or unexpected results occur, define who decides whether to suspend the workflow and how to return to the previous process.

A useful decision brief records the use case, chosen pattern, permitted data, excluded cases, required review, metrics, thresholds, and responsible person. It also notes what is still unknown. This makes it possible to compare alternatives or identify new use cases without starting the risk and operations questions from scratch.

Test before expanding

  1. 01Select a representative sample and add edge cases and out-of-scope requests.
  2. 02Define success and stop criteria before reviewing the results.
  3. 03Review outputs and errors by task, category, source, and consequence.
  4. 04Compare with the current process, including review time and escalations.
  5. 05Document failures, corrections, owners, and the conditions for expanding, maintaining, or stopping the test.
08

Decide based on evidence and keep review open

The decision is not simply whether to adopt AI or reject it. You might keep conventional automation, test classification for a single category, use search with reviewed drafts, or pause the initiative until the data and process are in better shape. If the main uncertainty is the quality of the documentation, a search test may be more informative than automating responses. If there is no way to verify an action or recover from an error, keep execution under human control.

To explore the options, separate three questions: which solution fits the task; how the possible patterns compare under the same conditions; and which other use cases might make sense later. This sequence helps you choose a specific scope, compare alternatives using consistent criteria, and identify opportunities without confusing them with cases that have already been validated.

The sources provided offer contextual guidance on enterprise AI, customer service, privacy, and agentic systems. They are not enough to determine which provider meets specific requirements, what an implementation will cost, how a service will retain data, or which current rules apply to a particular sector or country. Confirm these matters in primary documentation and with the responsible functions before sending data, contracting for a service, or automating a decision.

The practical final criterion is simple: start with a bounded task, retain human involvement where consequences require it, and expand only when a representative test shows that the outcome is acceptable, operable, and sustainable. If the evidence is insufficient, abstaining or keeping the current process may be the right decision.

Open questions

  • The supplied information about the sources is descriptive and partial; it does not include enough primary documentation to compare effectiveness, costs, or outcomes across technical patterns.
  • No current conditions are provided for data retention, security, pricing, or availability of specific services; these must be checked directly in primary documentation before use.
  • The note about the Spanish Data Protection Agency's agentic AI guide confirms its general relevance but does not support attributing specific design or authorization recommendations to it.
  • The available evidence does not determine which rules apply to a particular organization, sector, country, or request type; that assessment requires legal and operational context.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction