Ilustración editorial para Amazon Nova 2 Lite vs Claude Fable 5.1 para extraer datos de documentos: cómo plantear una comparación reproducible
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What decision this comparison addresses—and what it cannot conclude

The relevant decision is not which model is “better” in the abstract, but which one fits a specific document-extraction workflow. The use case considered here is converting invoices, short contracts, forms, and PDF documents containing tables into data that complies with a defined schema. In these workflows, an answer that appears correct is not enough: it must preserve relevant values, respect required types, distinguish an absence from an actual value, and leave enough traceability to correct errors.

The information provided describes Amazon Nova 2 Lite and Claude Fable 5.1 through their providers, but it does not include the results of a shared evaluation run on both. It is therefore not possible, based on these sources, to claim that one achieves higher accuracy, lower latency, or lower effective cost than the other for a particular document portfolio. Nor would it be rigorous to extrapolate Amazon comparisons against other models to the matchup proposed here.

The operational conclusion before testing is conditional. Nova 2 Lite can be evaluated as a multimodal model in the Nova 2 family through Amazon Bedrock. Claude Fable 5.1 can be evaluated through the access route and under the availability conditions documented by Anthropic. The selection will only be defensible after fixing an access route, a frozen corpus, a definition of correctness, a retry policy, and a consistent method for allocating auxiliary costs.

A useful comparison must separate three layers. The first is the capability declared by each provider: supported modalities, APIs, document handling, and available controls. The second is the result measured on a specific document set. The third is operational risk: incomplete data, defective OCR, non-equivalent documents, version changes, quotas, retention policies, and human review. Confusing these layers often produces a decision that looks quantitative but is difficult to audit.

02

Models, access routes, and asymmetries that must be fixed

Amazon documents Nova 2 Lite in Amazon Bedrock, including model identifiers, inference options, and possible regional or global deployment. Nova documentation also describes capabilities related to document understanding, PDFs, tables, and layouts. Those capabilities must be verified in the exact contracted modality, because a text-only test is not equivalent to a test that supplies images or documents to the model.

Anthropic identifies Claude Fable 5.1 in its documentation and product page, where it also communicates availability, pricing, and capabilities related to PDFs and tables. Before comparing the models, the team should record the exact identifier used, the date, the region or processing location where applicable, the API version, configured limits, output budget, and any reasoning or caching configuration. Without this record, a later result may not be reproducible.

There is a potential methodological asymmetry in structured output. Anthropic’s structured-output documentation describes a JSON Schema-based mechanism and warns that compatibility depends on platform and model. According to the supplied information, Fable 5.1 is not listed among Bedrock models compatible with structured outputs. This does not prove that Fable 5.1 cannot produce valid JSON through other routes; it does require documenting which mechanism was used on each side and avoiding the presentation of two different integrations as though they were identical.

There are two acceptable options, each with different implications. The first is to use each provider’s native interface and measure the system that can actually be deployed, while declaring platform differences. The second is to impose a common denominator: a prompt requiring JSON, external validation against the same schema, and an identical number of retries. The second reduces the effect of integration-specific advantages; the first better reflects the operational experience. The two should not be mixed without labeling them as separate experiments.

Comparability decisions before execution

VariableRecommended ruleRisk if it differs between models
Access routeRecord provider, API, region, and date for every runAttributing platform-caused differences to a model
Document inputProvide the same file and the same representation whenever possibleMeasuring prior OCR or conversion rather than extraction
Structured outputUse the same JSON Schema and a shared external validatorConfusing accepted formatting with semantic accuracy
ParametersFix temperature, output limit, and reasoning budget where applicableIntroducing configuration-driven variation
RetriesApply an identical, limited, and recorded policyHiding failures through additional attempts
CostInclude tokens, cache, files, retries, and auxiliary servicesUnderestimating cost per useful extraction
03

A reproducible design for structured extraction

The corpus should be frozen before results are observed and should represent the work intended for automation. A practical split includes clean digital text, simple tables, multi-page tables, scans, forms, documents with ambiguous fields, and incomplete documents. The original file, a non-sensitive identifier, document type, language, input channel, and known issues should be retained. If real data is used, the evaluation set must comply with applicable internal and contractual policies.

Every document needs ground truth independent of the model. For critical fields, two reviewers should annotate the expected value, legitimate absence, ambiguity, and the visual or textual evidence supporting it. Where they disagree, a third review or adjudication rule must resolve the case. The percentage of double-reviewed documents and unresolved disagreements must be published as part of the uncertainty, not hidden behind a single accuracy label.

The schema should represent both the data and its status. For example, a date can be valid, absent, illegible, ambiguous, or out of scope. Always forcing a string can turn an omission into an invention. It is useful to require evidence fields, a confidence level expressed under explicit rules, and an issues list; however, confidence declared by a model does not prove that it is well calibrated. It must be measured against observed errors.

The prompt must be identical in its semantic objective: extract only what is visible or explicitly inferable under a documented rule, return absence where evidence is missing, and do not complete plausible values. If one provider’s interface adds system instructions or mandatory formatting elements, those should be preserved, archived, and counted as part of the configuration. Any prior PDF conversion, OCR, image downscaling, or page splitting must also be recorded.

Repeated runs are necessary even when the configuration appears deterministic. At a minimum, every combination of model, document type, and configuration should be run three times. Results must report dispersion and, where sample size permits, uncertainty intervals. An isolated difference of a few fields should not become a purchasing recommendation if it falls within observed variation or comes from too few documents.

Minimum evaluation process

  1. 01Freeze the corpus, schema, normalizers, and acceptance criteria before the first run.
  2. 02Annotate ground truth and resolve disagreements between reviewers.
  3. 03Run every configuration at least three times using the same files and recorded parameters.
  4. 04Validate syntax and JSON Schema outside the model; then compare every field with the reference.
  5. 05Where applicable, apply the same limited retry to both systems and retain every attempt.
  6. 06Calculate metrics, full costs, dispersion, and errors by document category.
  7. 07Review a sample of failures to distinguish a model error, an OCR failure, an ambiguous rule, or an evaluator defect.
04

Metrics that matter: format, meaning, time, and cost

The valid JSON rate measures how many responses can be parsed without a repair that changes their content. It is necessary but insufficient. Perfectly valid JSON can contain the wrong supplier, an invented amount, or a shifted date. It should therefore be combined with field-level accuracy and document-level accuracy. The latter can be defined strictly: a document counts as correct only if all critical fields match the reference and no unsupported values appear.

Omissions and hallucinations must be counted separately. An omission is a field that should have been extracted but is missing or incorrectly marked as absent. A hallucination is a value presented as extracted without support in the document or contrary to the reference. In contracting, billing, or compliance, hallucinations can carry a higher operational cost than omissions, because a review can detect a blank more easily than a plausible but wrong datum. Weighting should reflect process risk, not a generic preference.

Uncertainty signaling should be evaluated as a classification task. If the model signals doubt in genuinely ambiguous or illegible cases, it provides value for routing to human review. If it declares high confidence in severe errors, that signal cannot support automated decisions. The report should include coverage: what fraction of the corpus can proceed without review under a fixed risk threshold; and conditional accuracy: how many accepted documents are actually correct.

Latency must be measured end to end, from the time the system starts the request until it receives a validated output or exhausts retries. Percentiles, not only averages, must be separated and measured by document size and complexity. File upload, conversion, OCR where applicable, inference, validation, and retry times must also be separated. Otherwise, an external bottleneck may be wrongly attributed to the model.

Useful cost is not the advertised token price. For each document, input, output, applicable cache reads or writes, file processing, retries, and auxiliary services are added together. Total cost is then divided by the number of documents that meet the definition of correctness. If human review is required, there are two separate measures: cost per correct automatic result and cost per final accepted result, including review work. Both are valid, but they answer different decisions.

Results table that should be published

MetricOperational definitionInterpretation
Valid JSONResponse passes the parser and schema without semantic repairMeasures technical integrability
Field-level accuracyCorrect fields divided by evaluable fieldsDetects localized failures
Document-level accuracyDocuments meeting all critical-field requirementsMeasures safe automation
Omission and hallucinationMissing errors versus unsupported valuesMakes it possible to weight operational harm
Safe coverageDocuments accepted without review under a fixed ruleMeasures how much work can be automated
End-to-end latencyTime until validated output or definitive failureMeasures experience and operational capacity
Cost per correct resultTotal cost divided by correct documentsConnects spending to actual usefulness
05

Results: what can be stated today and how to interpret a future test

No measurements from the proposed corpus have been provided for Amazon Nova 2 Lite or Claude Fable 5.1. Therefore, the results sections must remain without a performance verdict until the protocol is run and published. AWS documentation can justify including Nova 2 Lite in the evaluation based on its declared document-understanding capabilities and related modalities. Anthropic documentation can justify including Fable 5.1 based on the capabilities and conditions it communicates for PDFs and tables. Neither description replaces an observed rate of correct fields.

A favorable result for Nova 2 Lite on clean text documents would not demonstrate superiority on scans, complex tables, or ambiguous contracts. Likewise, a favorable result for Claude Fable 5.1 in a configuration using a structured-output feature would not demonstrate that the advantage comes from the model rather than the integration. The breakdown by document category is not an editorial detail: it is the basis for a team to apply the conclusion to its own portfolio.

AWS publishes an architecture example that combines Nova 2 Lite with another Claude model for scanned-document digitization. That material is useful as an illustration of a multi-stage architecture and of the need to allocate costs and tasks. It must not be used as comparative evidence between Nova 2 Lite and Claude Fable 5.1: the cited example uses another Claude model and a different document case.

A future test should also report recurrent failures, not only aggregates. Useful categories include: a table cell assigned to the wrong column, confusion between subtotal and total amount, incorrectly normalized date, wrong contracting party, an error caused by rotation or low resolution, a missing page, a non-visible value, and schema noncompliance. A set of anonymized and reviewable examples helps identify whether a high average hides an unacceptable failure type.

06

Privacy, retention, and conditions that can invalidate a metrics-only decision

The choice of access route affects both the evaluation and deployment. Amazon Bedrock documents data-retention policies and security and privacy controls. Anthropic separately documents retention conditions for its API, including options such as zero data retention where available under applicable conditions. These policies must be verified for the specific account, product, region, and contractual agreement; a general statement about a provider does not replace that verification.

Before uploading real documents, the team should classify the data, determine whether it contains personal, financial, contractual, or regulated information, and confirm who can access prompts, responses, files, and logs. It should also check whether the JSON schema, traceability metadata, and error examples contain sensitive information. Data minimization may matter more than a small latency or price difference.

Data residency, encryption, identity management, customer-managed keys, and network controls may determine which service is admissible. AWS describes enterprise controls for Bedrock, including security and isolation mechanisms. The evaluation should record which controls were enabled, because a laboratory configuration with broad permissions may not represent an architecture that can be approved for production.

Teams should not assume that data is not used for training, that it is deleted immediately, or that it is processed in a specific location without checking the current documentation and contract. Policies change and may depend on the access modality. If a compliance requirement is not confirmed, the decision should be “pending validation,” not a provisional technical recommendation presented as sufficient.

Privacy gate before the pilot

  1. 01Identify data categories, jurisdictions, and retention obligations.
  2. 02Confirm the access route, region, retention configuration, and applicable agreement.
  3. 03Review identity, logging, encryption, and file-access controls.
  4. 04Apply minimization, pseudonymization, or synthetic data where feasible.
  5. 05Approve a limited sample before expanding the corpus or moving to production.
  6. 06Document which statements depend on contract or configuration and require periodic review.
07

Scenario-based decision-making and the limits of this comparison

For a lowest-acceptable-cost scenario, the decision should be based on cost per correct document within a minimum accuracy threshold, not on the advertised unit price. For maximum accuracy, document-level accuracy on critical fields, hallucination rate, and uncertainty calibration should prevail. For lowest latency, teams should examine end-to-end percentiles under the expected load. For high criticality, the right option may be a combination of extraction, deterministic validation, and human review, even if a model achieves a strong aggregate average.

When both models fail a defined threshold for critical fields, the correct conclusion is that neither should automatically approve those documents in that configuration. Alternatives include improving input quality, separating OCR from extraction, splitting documents into pages or sections, narrowing the schema, adding validation rules, using risk-based routing, or retaining human review. Changing models without diagnosing the cause can move the error without solving it.

The update schedule should depend on material changes: a new model identifier, a pricing change, modified limits or modalities, updated data policies, a region change, a shift in the production corpus, or the arrival of a new document type. Every new measurement should preserve the previous version to avoid comparing incompatible results. The recommendation should explicitly expire if components change without being retested.

In summary, the sources establish that documented capabilities and controls exist that justify evaluating both options, but they do not quantify an advantage between them for reliable document extraction. A traceable decision requires an in-house, repeatable test. Until it is published, any preference should be treated as an integration, cost, or compliance hypothesis—not as a demonstrated quality result.

Decision rule by priority

PriorityTie-breaking metricNo-automation condition
CostLower cost per correct document after retriesDoes not meet the minimum accuracy threshold
AccuracyHigher document-level accuracy on critical fieldsHallucinations in high-impact fields
LatencyBetter end-to-end percentile with validated outputQueues or retries exceed the SLA
ComplianceRoute meeting verified controls and conditionsRetention, residency, or access is unconfirmed
High riskBest safe coverage with human reviewPoorly calibrated uncertainty or undetectable errors

Open questions

  • No execution results, corpus, sample size, test dates, effective regions, or parameters were provided to compare the performance of the two models.
  • The exact compatibility of Claude Fable 5.1 with structured-output mechanisms depends on the access route and must be confirmed in the chosen configuration.
  • Effective pricing, limits, regional availability, and retention conditions can change and require verification on the contracting date.
  • The percentage of ground truth reviewed by two people and the discrepancy-resolution policy for the proposed corpus are unknown.
  • The effects of OCR, PDF conversion, storage, and other auxiliary services cannot be attributed to the models without an equivalent experimental architecture.
08

Keep exploring

08

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction