The Unit of Decision Is the Audio Workflow, Not the Provider
Saying that a team “uses ElevenLabs” provides little useful information for operations, risk reviews, or procurement. A deployment may bring together text-to-speech synthesis, transcription, professional voice cloning, shared voices, Creative Studio projects, and conversational agents. Those capabilities are not equivalent in terms of configuration, permissions, or information persistence.
The useful unit for an inventory is each specific workflow. For example, a customer-support application may send text to a synthesis endpoint using a particular voice; an editorial team may work in a Studio project; and an agent may produce recordings and transcripts. Although all of these belong to the same provider and workspace, they should not be assumed to share a model, access controls, retention settings, or deletion mechanism.
The first deliverable that should be required before launch is a record for each workflow. It should connect the business purpose with the technical model identifier, the voice identifier where applicable, the product or processing channel, the environment, the identity that performs the operation, and the treatment of the data. This record does not, by itself, prove regulatory compliance or authorization to use a voice, but it makes questions verifiable that would otherwise remain scattered across code, dashboards, and commercial discussions.
Initial Inventory for an Audio Workflow
- 01Define the input, output, and purpose—for example, support text converted into audio for an outbound call.
- 02Record the exact product and channel: API, web interface, Creative Studio, or ElevenAgents; do not group them under a generic label.
- 03Note the configured model and identifier, the voice ID where applicable, the output format, and any parameters that may alter the result.
- 04Classify the voice as a shared or library voice, designed voice, instant clone, or professional clone; retain evidence of its origin and permission separately from the identifier itself.
- 05Assign a technical owner and a data owner, establish the expected retention regime, and document the testing and deletion procedure.
- 06Define the expected replacement and regression tests for a model retirement or change.
Model Layer: Identify the Technical Contract Being Invoked
The model documentation distinguishes text-to-speech, speech-to-text, and conversational voice capabilities, and exposes model identifiers together with properties such as languages, modality, and limits. The commercial name of a feature is not enough to reconstruct an integration: configuration should retain the identifier received by the call and the date or version of the application configuration.
This precaution matters especially when models are deprecated. The provider’s documentation identifies deprecated models and publishes replacements for certain cases, including eleven_turbo_v2_5, eleven_turbo_v2, and scribe_v1. A recommended replacement is a starting point for planning, not a guarantee of functional equivalence. A change may affect response times, behavior across languages, transcription format, consumption, downstream automations, or editorial expectations.
An integration should treat retirement as a normal part of its lifecycle. It is advisable to monitor notices received by the team, keep configuration explicit rather than relying on defaults, and run parallel tests on a representative sample. The assessment should cover both audio or transcription results and operational effects: errors, observed latency, limits, internal costs, and compatibility with systems that consume the output.
Decision Questions for the Model Layer
| Field | Evidence to retain | Decision it enables |
|---|---|---|
| Model identifier | Value configured in code, an environment variable, or version-controlled configuration | Determine which model generated an output and locate workflows affected by retirement |
| Modality | Applicable documentation and a recorded test of the specific workflow | Distinguish a real-time capability from a batch execution without inferring it from the name |
| Limits and languages | Documented properties for the model, separated from the results of the team’s own tests | Define input validation and contingency scenarios |
| Model status | Current status and the documented replacement where one exists | Plan a migration before the service becomes unavailable |
| Acceptance criterion | Internal metrics, test cases, and approval owner | Decide whether to promote, roll back, or stop the replacement |
Voice Layer: A Voice ID Does Not Establish Ownership or Permission
A voice asset should be treated as a resource separate from the model. The same model can be used with different voices, and a voice can be available to multiple workflows depending on the permissions granted. A record of a generation should therefore retain the voice ID used, but also a reference to its provenance record, declared ownership, authorized purpose, territory or term where relevant, and agreed withdrawal conditions.
Professional voice cloning deserves explicit separation. The documentation describes a workflow that includes uploading samples, subsequently using a voice ID, and verification before training. The owner of the voice must read and record a verification challenge. That verification is a technical condition of the described process; it does not replace the team’s assessment of the applicable authorization basis, agreed terms, an organization’s representation, or relevant usage restrictions.
In addition, the documentation indicates that professional voice cloning requires voice material to be retained in order to generate. This has a practical consequence: a no-retention policy should not be promised for a professional voice-cloning workflow without reviewing the specific product and its configuration. The decision to retain, delete, or withdraw a voice should account for samples, the trained asset, outputs already generated, and copies subject to a different regime.
Library or shared voices, designed voices, and instant clones also require classification, even if their technical requirements and evidence requirements are not identical to those of professional cloning. It is not prudent to infer the origin of a voice solely from its name, its resemblance to a person, or its availability in an account.
Channel Layer: API, Interface, Studio, and Agents Must Not Be Treated as Synonyms
The channel through which content enters may materially change which controls can be configured. An API call made by a service account, an action in the web interface, a Creative Studio project, and a conversation handled by ElevenAgents should be inventoried separately. The fact that a measure is available for one type of traffic does not allow it to be extended automatically to the others.
This is particularly important for Zero Retention Mode, which is available for Enterprise. The documentation defines eligible endpoints, excludes web-interface traffic, and lists ineligible products, including cloning and Studio. It also describes limitations related to support, deletion, and backups. Consequently, the name of the mode should not be turned into a broad statement such as “no data is retained.” The operationally correct wording is more specific: identify whether the endpoint, product, workspace, and traffic for the workflow fall within the documented scope, and retain evidence that it was enabled.
In ElevenAgents, transcript and recording retention is configured specifically. The documentation identifies two years as the default value and allows a number of days, unlimited retention, or scheduled deletion to be defined. These options require a decision about which artifacts the service needs, who may change the setting, and what happens to data already collected. They also do not permit inferences about the behavior of the API or other products.
Operational Security: Limit Access to Voices and Projects
Service accounts make it possible to separate integration identities from personal accounts. The documentation states that they start without access to resources and receive permissions through groups or direct sharing. This supports a least-privilege pattern, but the outcome depends on how the workspace is configured: creating a service account does not itself prevent excessive access if it is added to broad groups or resources are shared without review.
Shareable resources include voices, Studio projects, and Agents. The documentation describes viewer, editor, and admin roles, as well as authorized principals. When designing an integration, it is advisable to grant only the resource and level required for the workflow. A narration automation does not necessarily need access to every project, agent, or voice in a workspace.
Traceability should make it possible to link a generation to the technical identity that requested it, the workflow that authorized it, the model, the voice, the configuration, and the output destination. Some of these records may need to be kept internally even if the provider applies a restrictive retention configuration. Security and data minimization should therefore be designed together: record enough to investigate an incident without unnecessarily storing input content, audio, or personal data.
Access Review Before Launch
- 01Create a separate service account for each integration or risk domain where feasible.
- 02Verify that it has no initial resource access, then grant explicit access only to the voices, projects, or agents required.
- 03Choose the minimum role compatible with the operation and document who approves each sharing decision.
- 04Store keys in a secrets-management system, assign each key a technical owner, and define rotation and revocation.
- 05Test in a separate environment that the identity cannot view or modify unrelated resources.
- 06Review groups, direct shares, inactive accounts, and evidence of use periodically.
Migration and Retirement: Design an Exit Before Depending on a Model
Migrations should not begin on the day a model becomes unusable. The inventory should associate each model with the workflows that consume it, relevant parameters, test data, and an owner. Where a documented replacement exists, the team can open a controlled assessment; where none exists, it should treat the change as an architecture decision that may require redesigned validations or user experience.
A useful test compares the behavior of the complete integration, not only an isolated demonstration. For text-to-speech, it may include texts of different lengths, intended languages, acronyms, numbers, proper names, and network failures. For transcription, it may include noise, overlapping speakers where relevant, and domain vocabulary. In both cases, internal thresholds and human review should be defined where the impact warrants them.
Voice compatibility requires an additional check. A replacement model may accept the same technical identifier without the team being entitled to assume an equivalent experience. The plan should define which changes are acceptable, who approves them, and which condition requires rolling back or stopping deployment. The provider’s documentation on replacements informs planning, while test results are local evidence for the specific use case.
Minimum Conditions for Promoting a Migration
| Area | Control question | Exit evidence |
|---|---|---|
| Model | Does the workflow use the replacement through an explicit configuration? | Reviewed configuration and deployed version |
| Functional quality | Do representative cases meet the internal criterion? | Test results and the owner’s decision |
| Voice | Does the authorized voice work as expected in the new workflow? | Approved sample and voice ID record |
| Operations | Are errors, timings, and limits acceptable? | Test metrics and observability plan |
| Data | Does the new channel retain information in line with its record? | Verified configuration and deletion procedure |
| Rollback | Can it return to the previous configuration or stop safely? | Tested runbook and an available owner |
Final Adoption Matrix and Conditions for Not Launching
An adoption matrix turns a general evaluation into reviewable controls. Each row should represent a workflow, not an entire product. It is preferable to mark a cell as “pending confirmation” rather than fill it with an inference drawn from a demonstration or a configuration in another channel.
No-launch conditions should be explicit. They may include: the model identifier actually invoked is unknown; the voice’s origin or permission is unsupported; the workflow uses a channel whose retention regime has not been verified; the service account has unnecessary resources; no owner is assigned to deletion; or a model retirement has no tests and no contingency procedure. These are internal governance decisions, not requirements universally declared by the provider’s documentation.
Finally, it is useful to separate three levels of assertion in internal and external documents. First, facts documented by the provider, such as a model status, a mode’s eligibility, or a retention option. Second, configuration verified by the team itself. Third, risk analysis and acceptance decisions. This separation reduces the risk of presenting an available feature as if it were enabled, or a technical measure as if it alone resolved legal, contractual, or editorial obligations.
Minimum Workflow Record for a Production Review
| Dimension | Minimum data | Acceptable status |
|---|---|---|
| Purpose | Use case, affected users, and owner | Specific, approved purpose |
| Model | Identifier, status, relevant limits, and replacement | Configured and verified |
| Voice | Type, voice ID, authorization owner, and withdrawal rule | Evidence located and current |
| Channel | API, interface, Studio, or Agents; endpoint where applicable | Channel identified without extrapolation |
| Data | Applicable retention, activation, exclusions, and deletion owner | Configuration verified |
| Access | Service account, shared resources, and role | Least privilege reviewed |
| Migration | Tests, acceptance threshold, and rollback | Executable plan before the change |
Open questions
- This piece is based on provider documentation supplied for verification. It does not independently assess the actual behavior of a particular account, endpoint, or contract.
- The documentation described does not establish that a setting is enabled in a specific workspace; this must be verified in the team’s configuration and tests.
- Authorization to use a voice, as well as legal, employment, contractual, or sector-specific obligations, depends on the case and is not demonstrated solely by a technical verification process.
- Prices, contractual availability, processing regions, and service-level guarantees are not detailed here because the supplied sources do not establish them precisely.
- Retention in the customer’s own systems may differ from the provider’s retention and requires a separate inventory.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction