Ilustración editorial para De responder a ejecutar: cómo cambió la interfaz entre las personas y la IA
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

A history told through what people can delegate

The history of artificial intelligence is often told through its techniques, its results, or the waves of enthusiasm surrounding it. It can also be told through a more concrete question: what could a person ask a system to do, and which parts of the task could they leave to it? This perspective focuses on interaction and delegation, not only on the ability to generate a response.

A convincing conversation, an interface built into a product, and an action carried out by software are different things. In the first case, the system produces an answer that a person can use. In the second, dialogue is presented alongside a service’s features. In the third, the system can intervene in an external process if it has the tools and permissions to do so. A conversational-looking interface is not enough to conclude that it can complete a task from beginning to end.

The sources available here offer several examples, either current or presented as context, but do not document the launch and evolution of a shared historical sequence in sufficient detail. This article therefore compares forms of interaction and defines what can be claimed about each example. It does not propose a universal chronology or argue that every new form is necessarily more reliable or useful than the one before it.

02

Five questions for distinguishing interaction from capability

To compare products without conflating their differences, it helps to separate five dimensions. The interface describes how a person communicates with the system. The technique or mechanism indicates, as far as the documentation allows, what the software does with the request. Scope of action specifies whether the result is an answer, a recommendation, or an operation carried out elsewhere. Supervision covers the permissions and confirmations involved. Finally, the available evidence clarifies whether a claim comes from product documentation, external analysis, or an evaluated demonstration.

This separation also helps us use related terms precisely. A prompt is the input or instruction that guides a response. An agent usually describes a system organized to pursue a goal through multiple steps, although usage of the term varies. Tool-calling refers to a model requesting the use of an available tool; it does not automatically mean the tool runs, the operation succeeds, or the task is completed without supervision. These terms do not, on their own, describe a product’s actual performance.

A comparison framework

Applying the same questions to every case helps prevent interface differences from being turned into a score for progress.

DimensionPractical questionWhat it cannot establish on its own
InterfaceHow does the person submit a request and receive a response?That the system understands every phrasing or context.
MechanismDoes the documentation describe generation, information retrieval, rules, or tool use?That the complete internal mechanism is known or that it always works.
ActionDoes the system respond, make a recommendation, or change something outside the conversation?That a recommendation has been carried out.
SupervisionWhat permissions, confirmations, or human reviews are required?That the system is autonomous or safe in other contexts.
EvidenceIs the claim in the provider’s documentation, or is there an independent evaluation?That the capability has been measured in a comparable way.
03

Text dialogue: an answer is not the same as an action

Text dialogue makes the interaction take the form of a question and an answer. A person writes a request, and the system returns text. As journalistic context, BBC Mundo describes ChatGPT as capable of answering questions or generating content. That description helps characterize a general conversational use, but it is not enough to establish the specific mechanism behind each response, the limitations that were tested, or which features were available on a particular date.

The distinction matters because an answer can guide a task without carrying it out. If someone asks for help drafting a message, receiving a draft does not show that the system sent it. If someone asks how to make a reservation, receiving instructions does not mean that availability was checked or a reservation confirmed. These are illustrative situations, not claims about the features of a specific product.

It is also not correct to infer that an early text interaction necessarily relied on rules or scripts simply because it is presented as a chatbot. IBM offers a historical overview that mentions early conversational chatbots, but the information available here does not explain how a specific system worked. Without product documentation or a direct historical source, it is not possible to reconstruct rigorously which rules, databases, or techniques were involved in a particular case.

The comparative lesson is modest but useful: conversation is a way of accessing a system, not complete proof of its capabilities. Evaluating a system requires information about the task, acceptable responses, errors, and conditions of use. A fluent exchange is no substitute for that evidence.

04

Assistants inside products: conversation situated within a service

An assistant embedded in a product changes where the interaction takes place: a person can converse in the context of a particular service rather than in a standalone interface. This integration may reduce the steps needed to find a feature or ask for help. However, proximity to a product’s features does not prove that the assistant can access all of them, correctly interpret every request, or carry them out without intervention.

The Adobe documentation available among the sources is a user-interface guide for its AI Assistant. The Primo Research Assistant documentation identifies a generative tool intended for research tasks. These materials can be used to describe examples of assistants presented within products or services. The information provided does not establish the exact capabilities they had at launch or allow us to reconstruct a sequence of changes with verified dates.

A different case, described by Facephi, concerns how to design a decision interface for analysts working with AI recommendations. This points to a relevant product-design question: showing a recommendation is not the same as replacing the judgment of the person receiving it. But, according to the available description, that source does not document a historical evolution or by itself demonstrate the results of a specific interface.

In practice, when describing any embedded assistant, it is worth checking what information it can access, which functions are available, what result it presents, and who confirms an operation. If the sources document only the interface or the general purpose, those are the limits of the description. It is not rigorous to fill in the gaps by assuming the system can access a user’s data or every feature in the product.

A process for verifying an embedded feature

This process helps distinguish a documented interface from a demonstrated action.

  1. 01Identify the feature described in the documentation and its date or version, if available.
  2. 02Separate the advertised capability from the action observed being completed.
  3. 03Check which permissions, data, and confirmations the operation requires.
  4. 04Record what happens in response to an error, an ambiguous request, or incomplete information.
  5. 05Limit the conclusion to what the documentation or a reproducible test establishes.
05

Tools and actions: a boundary that requires evidence

When a system can request the use of tools, interaction can go beyond producing text. A tool might, for example, retrieve information or start an operation in another service. But several stages must not be conflated: the model may propose a tool call; the software may validate it; an external service may respond; and a person may have to authorize the result. The existence of one of these stages does not prove that all of them are completed or that the process is correct.

The term tool-calling names a technical possibility, not a guaranteed outcome. An agent may coordinate several steps, but the label does not establish that the system chooses the right steps, maintains context, recovers from failures, or knows when to stop. To support a claim that a product carries out a specific action, we need sources describing the operation and, where possible, a reproducible demonstration of its limits and requirements.

The sources gathered for this article do not provide a sufficiently detailed documented example of a system invoking a tool and completing an external action, nor do they describe permissions and confirmations for such a case. For that reason, no such capability is attributed to Primo Research Assistant, Adobe’s assistant, or any other product mentioned. The contrast with conversational assistants is used here to formulate what would need to be verified, not to claim that a new historical stage has already been proven.

06

What can—and cannot—be compared

A fair comparison would use the same task and equivalent criteria for each system. For example, it could examine whether a person can find information, draft a response, or complete an administrative task. But first we would have to select specific systems, fix their versions, and obtain enough documentation to understand the task, permissions, human involvement, and outcome. The available evidence does not let us reconstruct the same task across the examples mentioned.

For that reason, this article compares categories and limits of evidence, not performance results. BBC Mundo’s context on ChatGPT and IBM’s reference to conversational chatbots offer general background. Adobe’s and Primo’s guides document specific products from their providers’ perspectives. Facephi’s article offers context on presenting recommendations and preserving human judgment. These are sources of different types and scope; they should not be treated as if they were comparable trials.

This caution helps avoid three common errors. First, treating a feature description as independent verification of its performance. Second, attributing capabilities to a product that are not mentioned in its interface guide. Third, calling any reduction in steps “autonomy,” even if a person still makes the important decisions. Integration, access to tools, and the quality of results are separate variables.

What the available sources allow us to claim

The level of detail should match the type of evidence, without filling undocumented gaps with assumptions.

Documented exampleCautious claimInformation that is missing
ChatGPT in the context described by BBC MundoThe source presents it as capable of answering questions or generating content.A comparable evaluation of reliability, mechanisms, and external actions.
Early chatbots in IBM’s overviewThe source mentions them as part of the historical context of AI.Primary documentation about a particular system and its specific mechanisms.
Primo Research AssistantThe provider’s documentation identifies it as a generative tool for research tasks.Capabilities at launch, evolution, limitations, and independent testing.
Adobe AI AssistantThe source is a user-interface guide for the product.A feature timeline and evidence that specific actions are carried out.
The decision interface discussed by FacephiThe source addresses how recommendations are presented to analysts and the role of human judgment.A historical evaluation or comparable product results.
07

More ways to interact do not amount to measured progress

An interface can make it easier for a person to express a need; an embedded assistant can put help alongside a service’s features; a tool can let software act on another system. Each change expands or reorganizes the possibilities for interaction. None, in isolation, establishes a shared improvement in reliability, usefulness, or general autonomy.

To support a claim of measurable progress, we would need to specify the task being evaluated, what counts as completing it, which errors matter, what level of supervision is allowed, and which systems are being compared. We would also need to distinguish observed performance from the commercial or documentary description of a feature. Without those elements, a claim may refer to a more convenient interface or a broader scope of action, but not to a generally superior capability.

Uncertainties are not a flaw to hide; they define what readers can learn from the available documentation. In this case, we have examples of conversation, assistants presented within services, and design centered on recommendations, but no consistent basis for telling a sequence of product stages or comparing their results. The map is therefore one of questions and distinctions, not a ranking of systems.

08

The key question is what is delegated—and what evidence supports it

Telling the story of AI through interaction makes it possible to see important changes without assuming they form an inevitable ladder. Asking for an answer, consulting an assistant inside a product, and delegating an operation are different experiences. To find out how much capability has really changed, we need to trace specific features, their dates, permissions, confirmations, and behavior when something goes wrong.

With the evidence available, we can distinguish text dialogue, embedded assistance, and the conceptual possibility of using tools. We cannot claim that the examples cited represent three successive stages, that the same task was completed better in each one, or that the named systems carry out external actions. Maintaining that boundary makes the comparison more useful: it prevents a new interface from being mistaken for a demonstrated capability.

When a future claim says that a product can now “do” something, the most practical question remains: what exactly did it do, under what conditions, who confirmed it, and where is it documented? The answer makes it possible to distinguish what the system suggests from what it actually carries out—and an interface novelty from a measured improvement.

Open questions

  • The available sources do not establish Primo Research Assistant’s capabilities at launch or describe its evolution.
  • The Adobe guide identifies an interface but does not, by itself, demonstrate a feature timeline or the execution of actions.
  • IBM’s overview is a secondary source, and the information provided does not allow us to reconstruct the mechanisms of a specific historical chatbot.
  • The sources provided do not offer enough evidence to compare the same task across the three types of interaction.
  • No verifiable case is documented here of tool use completing an external action, including its permissions and confirmations.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction