Define the decision you are testing
Screening titles and abstracts means deciding which records from a literature search appear to meet the eligibility criteria and should move to the next stage. It does not mean confirming that a study is eligible: abstracts often do not contain enough information, and a final decision may require consulting the full text. The pilot proposed here therefore evaluates a limited task: classifying records to decide which ones need further review. It does not complete a systematic review or replace its protocol.
The practical question is not “Can AI screen studies?” but “Which method helps this team, with these criteria and this set of records, without increasing the risk of missing relevant studies?” A useful answer considers errors, consistency, traceability, and human working time. A tool producing a quick label does not show that the label is correct or that the net time saving is meaningful.
It helps to distinguish three approaches. In manual screening, people read records and record their decisions, for example in a spreadsheet. A specialized tool is designed for tasks related to literature reviews, but its specific functions must be checked in the product and version you intend to use. A general-purpose AI model can be prompted to classify text, but do not assume that it provides a screening workflow, a decision log, or controls suitable for a review. Configuration matters in every case.
Set criteria and test cases before comparing options
Before opening a tool, turn the review question into inclusion and exclusion criteria that can be applied to individual records. Specify the population, intervention or exposure, comparator, outcomes, study design, context, and any date or language limits that are genuinely part of the protocol. Not every element will matter for every review; the important thing is to state what information determines eligibility and what to do when it is missing.
Add instructions for ambiguous cases. For example: if the abstract does not let you confirm a criterion, should the record be retained for full-text review or excluded? What should happen when a study seems relevant but does not state its design? What if the title appears to meet the criteria but the abstract contradicts it? If the team does not resolve these questions in advance, two methods may appear to disagree when they are actually applying different or incomplete rules.
Prepare a set of records that includes clearly eligible, clearly ineligible, and uncertain examples. An institutional study-selection guide recommends piloting criteria with records from these three groups; this article adapts that guidance for the additional purpose of comparing alternatives. Preserve titles and abstracts as retrieved, except for transformations that are necessary and documented to protect data or normalize formatting. Do not select only easy examples: the pilot should reveal where the instructions fall short.
Establish a reference decision through independent human review and a procedure for resolving disagreements. The reference is not infallible truth; it is a decision documented in line with the protocol, with uncertainties made visible. If reviewers disagree, adjudicate the case, clarify the relevant rule, and record the change. Then apply that same rule to all three alternatives.
Minimum preparation for the test set
- 01Freeze a version of the review question and eligibility criteria.
- 02Select eligible, ineligible, and ambiguous records; document how you selected them.
- 03Ask two people to apply the criteria without consulting the tools’ labels.
- 04Resolve disagreements and retain the rationale and rule used.
- 05Separate records used to refine instructions from those reserved for testing them.
Compare the three approaches without assuming unverified capabilities
Manual screening applies the team’s criteria directly, but requires people to read and record every decision. A spreadsheet can help organize records and document labels if the team sets it up for that purpose; it does not, by itself, provide a methodological decision or resolve disagreements. To make the comparison fair, define the columns, permitted labels, and who may change a decision in advance.
Specialized tools are candidates for a pilot, not a guarantee of quality. Rayyan, Elicit, SciSpace, and scite.ai appear in the supplied sources as tools related to reviews or bibliographic analysis. That mention can guide an initial shortlist, but it does not confirm that each currently offers title-and-abstract screening, that it does so in a way that fits your protocol, or that the relevant functions are available under equivalent access conditions. Verify each task in the product and its current documentation before designing the test.
A general-purpose model may be useful for testing explicit instructions based on your criteria, but any result will depend on the text provided, the instructions, the configuration, and the response it generates. Do not assume that it will return the same classification in every run, cite a verifiable reason, or retain an adequate history. Check those properties directly and record the visible version, date, and available parameters. If the system changes between runs, that variation is part of the result, not a detail to conceal.
Apply the same eligibility information to every alternative. If a tool cannot export a decision, retain an explanation, or reproduce a configuration exactly, record that as an operational limitation. Do not try to make up for a missing function by assuming what the provider probably offers.
Initial comparison: what to check for each approach
| Approach | What to test | What not to assume | Minimum record |
|---|---|---|---|
| Manual with a spreadsheet | Human application of the criteria and the agreed annotation workflow. | That the spreadsheet controls consistency or resolves uncertain cases. | Decision, reviewer, rationale, and status for each record. |
| Specialized tool | Screening functions confirmed in the product and test account. | Current availability, price, export, explanations, or fit with the protocol. | Product and version, configuration, decisions, and verified export options. |
| General-purpose AI model | Classification of titles and abstracts using fixed instructions. | Repeatability, traceability, persistent access, or sufficient accuracy. | Visible model or version, full instructions, date, output, and human review. |
Design a comparable and reproducible test
Use the same records, criteria, and rules for resolving uncertainty with all three approaches. Write one clear set of instructions; if a tool requires a different format, preserve the meaning and document the adaptation. Do not improve one option’s instructions after seeing its errors without repeating the tests for the others: that would favor the alternative that received more adjustment.
Separate test development from evaluation. You can use an initial group to clarify instructions and identify formatting problems; reserve another group to test the results with the instructions frozen. If the available corpus is small, report that limitation and avoid presenting a few cases as conclusive evidence of general performance. The purpose of the pilot is to obtain evidence relevant to your team’s workflow, not to declare a universal winner.
Where the workflow permits, ensure that the people who established the reference decision do not see the system’s labels first. Then compare each output with that reference and manually review every exclusion proposed by an AI in the evaluation sample. This control is especially important because an incorrect exclusion could remove a study that would have been relevant. Retain disagreements too: they may indicate that a criterion needs to be clearer or that a case requires adjudication.
Test repeatability when a method may produce different outputs. Repeat part of the test set with the same configuration and compare decisions. For a tool whose configuration cannot be fixed, record what controls it actually offers. Lack of control does not, by itself, prove a classification is wrong, but it does limit how well the pilot can be reproduced and audited.
Measure errors, consistency, and workload
Do not reduce the evaluation to a single accuracy figure. Count false exclusions: records the alternative excludes but the reference retains for full-text review. Also count false inclusions: records the alternative keeps even though the reference excludes them. Report uncertain cases and disagreements that required review or adjudication. Present the counts alongside the size of the test set and explain how the reference was established.
Errors do not all have the same practical impact. For a team that prioritizes not missing potentially relevant studies, a false exclusion may be more concerning than reviewing some additional records. State that priority as a decision criterion before viewing the results; do not select it afterward to justify the preferred option. If the protocol requires independent decisions by multiple people, the pilot should not replace that requirement without an approved methodological justification.
Record time separately for preparation, import, execution, review of outputs, resolution of ambiguous cases, and export or data cleanup. Compare total human work, not just the time a tool takes to produce labels. Also measure whether the team can reconstruct what happened to a record: who made the decision, which criterion they used, what information they relied on, and whether the decision was later changed.
An alternative that reduces initial reading time may require more verification work. Conversely, a manual workflow may be simpler for a small corpus or criteria that are difficult to formalize. The pilot should show these costs in the context of the actual team. Do not present a difference observed in a limited test as a universal property of a product.
Pilot evaluation framework
| Measure | Question it answers | How to record it |
|---|---|---|
| False exclusions | How many reference records for full-text review were excluded? | Count and list of identifiers for review. |
| False inclusions | How many records excluded by the reference were retained? | Count, criteria involved, and estimated additional work. |
| Uncertainty and disagreement | Which cases could not be resolved without human intervention? | Count, reason, adjudication, and changes to criteria. |
| Consistency | Did labels stay the same when the test was run again? | Number and type of decisions that changed under documented conditions. |
| Human work and traceability | How much effort was needed to validate and reconstruct decisions? | Time by phase and available audit fields. |
Consider limitations, data, and decisions that need oversight
An incomplete title or abstract may conceal decisive information. A system cannot reliably resolve a criterion that is absent from the record without consulting another source; asking it to fill the gap may produce an inference that sounds confident but is not supported by the text. Define an uncertainty option and use it. When eligibility depends on methodological details missing from the abstract, retain the record for the full-text stage.
Criteria may also be poorly operationalized. Phrases such as “relevant population” or “adequate study quality” are not enough unless the team has stated what textual evidence satisfies them. If people and systems interpret a rule differently, review the rule before attributing the problem solely to the tool. After any change, document the criteria version and reapply the reserved test set using a consistent procedure.
Before uploading records, review the applicable terms for access, data processing, storage, retention, and reuse for the service you plan to test. The supplied sources do not provide current access, pricing, or data-processing conditions for the candidate products; these details need to be checked directly. If the materials contain sensitive information or institutional restrictions apply, consult your organization’s rules and do not send data until you have authorization.
Keep approval of the criteria, resolution of uncertain cases, and the final eligibility decision under the team’s responsibility. A university guide to AI in evidence synthesis discusses using these tools alongside expert judgment; it is guidance, not evidence that a particular function is safe or effective. Transparency means recording what was automated, what was reviewed, what changed, and what limitations the test had.
Choose with a decision matrix, not a promise
At the end, compare results with the priorities agreed before the pilot. If the team cannot tolerate false exclusions, rule out any configuration that produces them without a review process capable of detecting them. If traceability is essential, require a workflow that can retrieve labels and explain any changes. If saving time is the goal, include supervision time, not just the tool’s execution time.
Keep screening manual when the set is manageable for the team, the criteria require contextual interpretation, or no automated alternative meets the minimum requirements for safety and record-keeping. Consider a specialized tool if its relevant functions have been verified, the pilot with your own records is acceptable, and the workflow allows results to be exported and audited. Test a general-purpose model only with controlled instructions, authorized data, and enough review to prevent its outputs from becoming unchecked exclusions.
If the result is promising, integrate the tool as an explicit change to the procedure: document its role, the versions used, human checks, disagreement handling, and how errors will be corrected. Reevaluate the workflow if the criteria, product, model, or configuration changes. If the evidence is insufficient, report the uncertainty and extend the pilot rather than claiming that one option is superior.
To follow Inferama’s editorial journey, this guide can link from the Discover, Compare, and Learn pages. That navigation does not change the methodological standard: the decision should be justified with pilot records and the needs of the protocol, not a marketing description or an unverified feature list.
Practical matrix for deciding what to do next
| Observed result | Prudent decision | Condition before proceeding |
|---|---|---|
| Unacceptable exclusions or exclusions that are hard to detect | Keep decisions with human reviewers; do not automate exclusions. | Review criteria and failed cases; repeat a controlled test. |
| Errors are controllable, traceability is adequate, and total work is lower | Consider limited use as screening support. | Approve the procedure, record human intervention, and monitor uncertain cases. |
| Product functions or conditions remain unverified | Do not compare as if the capability were confirmed. | Check the workflow, export, access, and data processing directly. |
| Variable results, a small corpus, or a disputed reference | Report uncertainty and extend or redesign the pilot. | Separate instruction tuning from evaluation; document adjudications. |
Open questions
- The supplied sources do not verify the current title-and-abstract screening functions of Rayyan, Elicit, SciSpace, or scite.ai.
- Current access conditions, prices, usage limits, and data-processing policies for the candidate products have not been provided.
- There are no independent comparative results or results from an in-house pilot that would support a general claim that one approach is more accurate or faster.
- The quality of the reference decision depends on the criteria and adjudication procedure established by each team.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction