Ilustración editorial para AEGIS: qué mide este benchmark de imágenes académicas manipuladas y por qué detección, explicación y localización no son el mismo resultado
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

AEGIS is not a test of scientific authenticity

AEGIS is a benchmark for evaluating forensic analysis of academic images generated or manipulated with AI. Its significance is not limited to asking whether a model recognizes synthetic content: it seeks to separate a classification decision, an explanation of the clues, and the spatial localization of a potential alteration. That separation matters because the three outputs answer different questions and may fail independently.

From a scientific-integrity perspective, an AEGIS result should be understood as a measurement under a defined protocol, not as certification that a figure is authentic or fraudulent. A high score may indicate that a system performed well on the benchmark’s examples, formats, annotations, and evaluation rules. It does not by itself establish that the system will behave the same way on an unseen image, a figure with different compression, a legitimately edited image with inadequate documentation, or a possible real case of misconduct.

This also defines its scope relative to generic synthetic-content detection benchmarks. An academic image often follows specific visual and semantic conventions: composite panels, scales, annotations, microscopy, charts, or other subtypes documented by the dataset. Context can provide useful signals, but it can also create shortcuts. A model may learn correlations with generation or editing styles represented in the test set without acquiring a general forensic capability.

The useful comparison is therefore not, in the abstract, “which system has the highest number?” It is: “Which task did it solve, with which input, on which split, under which correctness criterion, and with which stated limitations?” Inferama’s benchmark, safety, and glossary routes can help frame this distinction between a controlled evaluation and an applied decision.

02

Three tasks, three kinds of claim

Binary detection asks a bounded question: under the benchmark’s operational definition, should an image be classified as real or as generated or manipulated? Its output is a label and may be accompanied by a score or probability. It can be useful for sorting cases that warrant review, but it does not necessarily identify the editing mechanism or show the visual evidence that led to the decision.

Reasoning about clues asks the system to state why it suspects an alteration. Under the reference format and evaluator specified by AEGIS, the response must relate to observable cues or to the represented forgery strategy. This task is not validated merely because the binary label is correct. A system may get the class right through a spurious correlation while also providing an explanation that is vague, inconsistent with the image, or produced after the decision rather than grounding it.

Localization requires identifying where the alteration is, usually through a region, mask, or other spatial representation that can be compared with a reference annotation. This is a different requirement from describing an anomaly in natural language. An explanation may mention a panel or visual element without delimiting it precisely enough; conversely, a plausible region does not demonstrate that the model correctly explained the nature of the manipulation.

These differences have a practical consequence: performance on one task cannot validly substitute for performance on another. Detection accuracy is not explanation accuracy; textual agreement is not a correct mask; and strong spatial overlap does not by itself establish that the final classification is reliable. Any result table should retain its task, metric, and protocol columns rather than compressing them into a single ranking of models.

What each output allows you to claim

Evaluated outputQuestion it answersSupported claimClaim not automatically supported
DetectionDoes the image match the class defined by the benchmark?The system classified examples under the protocol.That it identified the manipulated area or editing mechanism.
ReasoningDoes the justification match the reference criterion?The system produced an explanation evaluated under that format.That its decision is causally based on that explanation.
LocalizationDoes the predicted region match the annotation?The system delineated the alteration under the spatial criterion used.That it can determine intent, authorship, or real fraud.
03

How to read metrics without treating them as interchangeable

Classification metrics such as accuracy summarize how many decisions match the dataset labels. Yet aggregate accuracy can conceal class-level differences. When the protocol reports class-specific metrics, they help determine whether performance is concentrated in a dominant class or whether the system handles real and altered images unevenly. Interpreting any percentage also requires the split size, class distribution, and rules for invalid or ambiguous answers.

The usual spatial metric in segmentation or localization tasks is intersection over union, or IoU. It compares the overlap between the predicted region and the annotated region. A low IoU may reveal an overly broad region, a displaced location, or a prediction that captures only part of the annotated area. But its interpretation depends on the annotation type, whether boxes or masks are evaluated, the thresholds used, and the handling of multiple regions or images with no alteration.

Reasoning calls for even more caution in interpretation. Evaluation may depend on reference answers, semantic matching criteria, or an automated procedure. The result indicates conformity with that procedure; it does not independently establish that an explanation is a causal reconstruction of the editing process. Before comparing results, check whether systems received the same image, the same instructions, the same opportunity to use external retrieval, and the same required output format.

A metric is not less valuable because it is limited. The problem is assigning it a meaning that the protocol does not measure. The dataset distributed by the project includes fields for task, reference answer, category, subtype, forgery strategy, and generative model. Those fields make it possible to disaggregate some findings, but a global figure remains a summary rather than a complete diagnosis.

A process for reading a published score

  1. 01Identify the task: detection, reasoning, or localization.
  2. 02Record the exact metric, the preferred direction of change, and the calculation protocol.
  3. 03Check the split, class distribution, and treatment of invalid cases.
  4. 04Look for results by category, subtype, forgery strategy, and generative family where available.
  5. 05Verify the input, prompt, threshold, and external resources received by the system.
  6. 06Limit the conclusion to what the task measures; do not extend it to authenticity or real fraud.
04

What the dataset contains, and why composition matters

The dataset card published by BUPT Reasoning Lab states that there are 20,571 rows and that the dataset is licensed under CC BY 4.0. Its described structure includes task information, reference answers, category, subtype, forgery strategy, and generative model. The data repository permits inspection of examples and associated resources, while an identified repository snapshot lists image directories and JSON files for real images and four forgery strategies.

This traceability is useful, but it does not replace an audit of the specific split used for a result. The number of rows should not automatically be interpreted as the number of visually independent images: the same image, or a related image, may appear in more than one task or have associated metadata. The relevant unit of analysis is the one established by the evaluation protocol and its separations between training, development, and test data.

Academic categories and subtypes are central to the difficulty of the test. A system may perform unevenly depending on visual conventions, level of detail, or the kind of signal available in each subtype. Likewise, different forgery strategies may leave different artifacts. If results are aggregated without breakdowns, they cannot reveal whether performance comes from a broadly distributed ability or from particularly distinguishable cases.

The presence of generative models in the metadata makes it possible to study generalization across generator families, provided that the protocol explicitly separates training and test conditions. That separation should not be assumed without documentation. A result is more informative when it states whether it evaluates images from seen or unseen generators, keeps an editing strategy out of the tuning phase, and prevents duplicates, near variants, or metadata from linking examples across splits.

05

What is actually being compared across models

AEGIS may bring together systems with different capabilities: general-purpose multimodal models, specialized forensic detectors, and unified approaches that produce multiple outputs. Their presence in the same table does not mean their conditions are identical. A specialized detector may take an image and return a score; a multimodal model may require instructions, generate free-form text, and depend on how that text is converted into an evaluable label or region.

Comparability requires disclosure of the exact model version, inference configuration, prompts, number of attempts, handling of large images, preprocessing transformations, decision thresholds, and resource budget. If retrieval or additional information is used, that must also be stated. An apparently small change—for example, a different rule for interpreting a textual answer—can alter a metric even when the underlying visual capability has not changed.

Results for a model such as Amazon Nova 2 Lite are interpretable against AEGIS only if a documented configuration exists for that evaluation. Membership in a multimodal model family does not establish a score, a localization capability, or suitability for scientific integrity work. The same applies to any product or specialized detector: a performance claim requires the protocol that produced it.

The official AEGIS repository documents benchmark execution and links JSON data, reference answers, images, masks, and optional retrieval resources. This architecture makes it possible to distinguish between what the evaluator calculates and what a system may use. However, anyone reproducing a result must record which resources were enabled and which materials remained inaccessible to the evaluated model.

Conditions that must match before comparing two scores

ElementWhat must be disclosedRisk if it changes
Dataset and revisionSplit, snapshot, and applied filtersDifferent examples are being compared.
SystemModel, version, tuning, and configurationA commercial name does not identify the evaluated behavior.
InputResolution, preprocessing, text, and auxiliary imagesOne variant may receive more visual or contextual signal.
InferencePrompt, temperature, retries, and thresholdDecision rules alter the metric.
EvaluationEvaluator, output format, and scoring ruleThe same output can be converted into different results.
06

What published results do and do not allow you to conclude

The AEGIS paper presents the benchmark, its tasks, dataset construction, metrics, and reference evaluations. That paper is the appropriate source for attributing results to the baselines studied by the authors. Critical reading should nevertheless retain the paper’s granularity: an aggregate average describes behavior under a particular mixture of examples, not a uniform guarantee across every category, subtype, or forgery strategy.

Diagnostic results are those that show where performance changes: differences between tasks, categories, subtypes, strategies, or generation families, where the study reports them. These breakdowns can reveal that detection appears more robust than localization, or that particular cases dominate an average. The cautious conclusion is conditional: under the evaluated configuration, the system performed differently in those groups. It does not authorize extrapolation of that pattern to external datasets without a new evaluation.

Nor is it correct to turn a low localization score into proof of an absolute inability to detect manipulation, or a high classification score into evidence of explainability. Each result identifies a distinct type of potential error. For an organization reviewing figures, this information may help design a prioritization and human-review workflow; it should not automate a sanction.

AEGIS does not by itself provide an estimate of fraud prevalence in the literature, a false-positive rate in a journal’s operational setting, or legal or institutional validation of an allegation. Those questions require samples representative of the use context, independent review criteria, appeal procedures, and prospective evaluation. They also require distinguishing deceptive alterations from legitimate, properly disclosed transformations.

07

Methodological risks: distribution, contamination, and calibration

The gap between simulated forgeries and real cases is a fundamental limitation. The strategies incorporated into the benchmark make controlled annotations and tasks possible, but real manipulations may be more heterogeneous, degraded by publication pipelines, or combined with procedures that are not represented. At the same time, authentic images in real use may include legitimate crops, contrast adjustments, compression, or panel composition that partially resemble editing signals.

Leakage or contamination can take several forms. A model may have encountered images, associated text, templates, or nearby resources during training; a team may repeatedly tune prompts on test items; or related examples may cross splits. Publishing materials helps reproducibility, but it makes it essential to document which data were public, which elements were held back, and how the evaluation was protected against iterative tuning.

Calibration matters when a score becomes an action. A threshold selected to maximize a benchmark metric may be unsuitable in a setting where manipulations are rare and the cost of a false positive is high. A deployed tool should report curves and errors relevant to its intended scenario, as well as criteria for escalation to human review. Without that information, isolated accuracy says little about operational utility.

Finally, transfer across domains must be demonstrated rather than assumed. A system evaluated on AEGIS categories may behave differently with new disciplines, instruments, annotation languages, journal formats, or generators. External evaluation, clearly separated from development resources, is the appropriate evidence for a claim of generalization.

08

A checklist before using an AEGIS score

Before reproducing a score, buying a tool, or including a result in a report, request evidence that makes the comparison reconstructible. The first question is which version of the dataset and evaluator was used. The second is what information the model received. The third is what decision is intended on the basis of the output. These questions connect technical design to institutional risks.

It is also useful to request the breakdowns that correspond to the decision at hand. If the objective is to prioritize images for review, false positives, calibration, and stability across subtypes are especially relevant. If a region is to be flagged, localization metrics and examples of errors should be examined. If an explanation is being assessed, review the criterion that defines an acceptable response, not only the fluency of the generated text.

AEGIS is most useful as a diagnostic instrument when it is reported together with its conditions and limits. It can then reveal what kind of evidence a system provides and what additional checks it needs. Treating it as an authenticity certification would erase precisely the distinctions between detection, reasoning, and localization that the benchmark is designed to measure.

Minimum questions for an evaluation or procurement decision

  1. 01Which dataset revision, files, and split were evaluated?
  2. 02What is the operational definition of each task, and how is it scored?
  3. 03Which model, version, prompt, preprocessing, threshold, and external resources were used?
  4. 04Are results available by category, subtype, strategy, and generative family?
  5. 05How were test-set tuning, contamination, and leakage across related examples prevented?
  6. 06Which false positives and false negatives are foreseeable in the intended use context?
  7. 07Is there external validation on academic images separate from the benchmark?
  8. 08Which human review, right of response, and primary evidence will be required before an adverse decision?

Open questions

  • The provided sources document AEGIS’s structure and evaluation, but this synthesis does not provide an independent reproduction of the benchmark or prospective validation in real scientific-integrity workflows.
  • The number of independent images should not be inferred from the stated 20,571 rows alone; the exact evaluation unit depends on the documentation and split used.
  • Applicability to disciplines, generators, publication formats, and editing practices not represented in the benchmark must be demonstrated through separate evaluations.
  • Any comparison involving a particular model requires a published AEGIS configuration; the model name alone does not establish a result.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction