The Short Answer
Generative AI can help students start systematic reviews faster by drafting candidate questions, search strings, and criteria—but the same speed can lead to over-delegation, superficial validation, and losing focus on conducting an auditable SLR. The new classroom experience report (10 doctoral students) explicitly observed these patterns across groups using ChatGPT support.
Practically, use AI to generate alternatives for early SLR steps, then slow down at checkpoints: verify alignment to scope, validate outputs critically, and document key decisions so your review remains traceable and defensible. This turns “tool use” into “method practice” rather than accepting AI text as the review.
The key caveat is that LLM outputs can feel persuasive while being wrong or misaligned, so without human supervision and decision records, you risk building a review that looks systematic but isn’t reliably evidence-based.
On this page
- Introduction
- Why This Matters
- Main Content Sections
- What the Classroom Study Actually Looked Like (and What the Researchers Were Watching)
- The Big Pattern: AI as a Barrier-Buster… and a Delegation Magnet
- The Manual-vs-AI Comparison Wasn’t Clean—And That’s an Important Lesson
- The Most Surprising Pedagogical Failure Modes (Beyond “AI hallucinations”)
- How to Use Generative AI for SLR Education Without Losing Methodological Rigor
- Key Takeaways
Smart Ways (and Risky Ways) to Use AI for Systematic Reviews—What a New Study Reveals
Introduction
If you’ve ever tried to teach (or learn) how to run a Systematic Literature Review (SLR), you know the hard part isn’t the theory—it’s the messy, chained decisions in the real workflow: shaping the research questions, building search strings, deciding inclusion/exclusion criteria, screening papers, and extracting evidence. New research from the original paper looks at exactly how generative AI shows up in that process—when doctoral students practice it in a real graduate course rather than just reading about it.
This experience report comes from a graduate Software Engineering class with 10 doctoral students, split into three groups, who piloted systematic reviews with and without generative AI support. The researchers didn’t just check the final outputs. They watched what students did during the session, collected artifacts, and analyzed conversation threads with ChatGPT assistants to reconstruct how students actually “appropriated” the technology—where it helped, where it derailed them, and how teacher mediation mattered.
The punchline is both encouraging and a little cautionary: LLMs can lower the initial barrier to starting an SLR and speed up producing candidate questions, search strings, and criteria. But the same speed can create over-delegation, shallow validation, operational headaches, and—most importantly—a shift in focus from “conducting an SLR” to “using the tool.” In other words: AI can make students fluent faster, but it can also make them trust faster than they should.
Why This Matters
This matters right now because many instructors (and research teams) are under pressure to “use AI,” often without clarity on how to keep systematic methods auditable. Systematic reviews are supposed to be traceable and justifiable—yet LLMs produce outputs that can feel confident even when they’re wrong or misaligned with your review scope. That creates a new pedagogical problem: students aren’t only learning the review method, they’re also learning how to supervise a reasoning partner that may not have your evidence standards.
A scenario you can apply today: imagine a small research group or thesis cohort that needs a rapid SLR for a practical project (e.g., “Do tools exist that detect bot-driven attacks?”). If they use an LLM to generate the research questions and inclusion criteria, they may finish the first draft quickly. But the study’s findings suggest that without structured decision records and reflection checkpoints, the team can accidentally optimize for convenience (or for what “looks systematic”) instead of coverage, validity, and transparency. That’s not just an academic risk—it’s the difference between a review you can defend and one you can only hope is right.
How does this build on earlier AI research? Prior work has explored LLMs across stages like question formulation, search strategy development, selection, and extraction. This paper doesn’t argue that LLMs should replace humans; instead, it zooms in on education-as-practice: what students do when the model is available, what they delegate, and where they struggle. It complements the “can LLMs help?” conversation by answering a more useful question for instructors: “When students are still learning the method, what patterns of AI use actually emerge—and what safeguards are needed?”
Main Content Sections
What the Classroom Study Actually Looked Like (and What the Researchers Were Watching)
The study is an experience report, not a controlled experiment. That distinction matters: the goal wasn’t to prove AI makes SLRs better. The goal was to document the experience and reconstruct decision-making patterns in a realistic teaching environment.
Here’s the setup:
- Participants: 10 doctoral students in Software Engineering
- Grouping: three groups
- Design: pilot systematic reviews with and without generative AI support
- Time format: a single-day classroom session (morning 8:00–12:00, afternoon 14:00–17:00)
- Topics chosen by the groups:
1. AI in security—detecting bot-driven attacks
2. AI in human resource management
3. Generative AI in education and its ethical implications
To understand how AI was used, the researchers combined:
1. Classroom observations (field notes)
2. Artifacts produced by students (e.g., research questions, search strings, criteria, extraction tables)
3. Interaction threads from ChatGPT assistants (used to reconstruct delegation, validation behavior, conflicts, and teacher mediation)
The focus: process over product
Even when the study wasn’t perfectly comparable across “manual vs AI” phases, it still generated something valuable: evidence about how students interacted with the tool during planning, search, screening, and initial extraction—the stages where decision-making is most visible.
That makes this report unusually practical for instructors: it’s about conduct, not just about outcomes.
The Big Pattern: AI as a Barrier-Buster… and a Delegation Magnet
One of the clearest positives was that LLM support helped students get started and iterate faster. But the same effect introduced a failure mode.
What worked well: early structure and faster iteration
In the planning stage and in building search strings and criteria, AI did three useful things:
- Reduced initial barriers: students could move from broad “state of the art” ideas to more operational research questions.
- Accelerated alternatives: instead of a single question or string, groups could request variations quickly.
- Made issues explicit: when AI evaluated a research question or search string, it often returned strengths and problems, giving the students concrete material to discuss.
This is consistent with why LLMs feel helpful: they lower friction. They can draft a structure when humans are stuck in “what exactly should this become?” mode.
What went wrong: AI “help” turning into AI “ownership”
The researchers saw a major risk: students sometimes shifted from using AI as a supportive reviewer to letting it become the decision-maker.
The paper describes AI roles that emerged in practice:
- Refinement assistant (good alignment with learning goals): students ask AI to review or improve something they started.
- Alternative generator (formative, but needs careful validation): AI proposes multiple options; the risk is accepting one without enough justification.
- Decision executor (highest risk): AI generates core protocol components—research questions, search strings, inclusion/exclusion criteria, and extraction structures—without students fully articulating why those decisions are valid.
That third role is where systematic rigor is most threatened, because selection criteria and search strings define what evidence you’re actually considering. If students delegate those choices, the review may look structured while quietly losing methodological accountability.
To make the comparison visible, here’s the paper’s emergent trade-off pattern:
| AI use pattern (as observed) | What students tended to do | Why it’s helpful | Main risk the study found |
|---|---|---|---|
Refinement assistant |
Ask AI to comment, review, or improve student drafts | Keeps students as owners of final decisions | Less critical risk; encourages reflection |
Alternative generator |
Request multiple candidate questions/strings/criteria | Speeds up exploration and makes options concrete | Students may accept outputs without checking scope + validity |
Decision executor |
Let AI generate/replace key protocol artifacts | Fast path to “looking systematic” | Over-delegation, weak methodological justification, superficial validation |
(Those themes are drawn directly from the reported observed patterns in the interaction threads and artifacts.)
The Manual-vs-AI Comparison Wasn’t Clean—And That’s an Important Lesson
The researchers initially planned a relatively tidy “manual first, then AI-assisted” comparison. In reality, classroom time pressure and practical constraints made the separation messy.
The paper explicitly notes that:
- The separation between stages and modes was less rigid than planned.
- Planning/search/screening had more consistent evidence for analysis.
- Extraction and synthesis were more heterogeneous across groups, producing fewer comparable artifacts.
Also, there were operational deviations:
- Some groups encountered tool limitations (like file/spreadsheet handling).
- In at least one case, operational issues led to the use of personal accounts / other AI instances. Because consent only covered institutional assistants, those interactions weren’t captured—so the researchers treat this as a lesson learned for infrastructure planning in future replications.
Why this matters for readers designing their own activities
If you try to run an SLR lab with AI, don’t assume the workflow will behave like a controlled experiment. In practice, students will:
- jump to AI when they’re stuck,
- mix manual and AI steps as needed,
- spend time figuring out how to get a structured output rather than debating methodological meaning.
So the practical takeaway isn’t “don’t do comparisons.” It’s “if you want reliable comparisons, build stronger checkpoints and recording formats so the process stays visible.”
The Most Surprising Pedagogical Failure Modes (Beyond “AI hallucinations”)
A common worry about LLMs is hallucination. This paper’s classroom observations highlight something broader: even when AI responses are plausible, students can still learn the wrong lesson.
Here are the failure modes the study observed:
1) Insufficiently critical acceptance
In some threads, the interaction followed a “request → response → move on” rhythm. Students didn’t always revisit implications. The paper describes cases where AI recommendations were treated as validations rather than prompts for justification—meaning students might produce artifacts that look correct but aren’t critically defended.
2) Methodological decisions guided by operational convenience
Students used AI to “make the activity feasible” in a limited time window. The risk: operational convenience can start driving selection criteria, which should instead be driven by review scope and evidence validity.
A concrete example described in the report: students asked AI for additional exclusion criteria to simplify screening, but they didn’t explicitly discuss risks like reduced coverage or potential selection bias.
3) Screening output became too compact—reasoning disappeared
To make screening efficient, groups used simplified categories like included/excluded/uncertain. The side effect: it reduced opportunities to inspect and discuss why each decision was made. In an educational context, that matters because justification is part of what students are supposed to learn.
4) “Tool time” expanded, shrinking “review time”
Although AI accelerated early artifact generation, the paper also reports time being consumed by:
- prompt tweaking,
- output formatting,
- file processing workarounds,
- consistency fixes and validations.
This yields a messy but realistic conclusion: AI speeds up drafting, but it also creates new supervision tasks.
5) Focus drift: SLR → “using the tool”
One of the study’s most important findings is the shift in focus. Students may start treating the LLM as the central work engine rather than the support mechanism. That’s a learning design issue, not a student capability issue.
How to Use Generative AI for SLR Education Without Losing Methodological Rigor
The paper closes with targeted adjustments for future iterations—and these are genuinely useful if you’re an instructor, a lab lead, or a curriculum designer.
Recommended changes (from the study’s lessons learned)
Split across more than one meeting
The single-day format created pressure for shortcuts, increasing excessive delegation risk.Add a short AI literacy training before the main activity
Students need practice with:- what LLMs can/can’t do in structured workflows,
- response validation,
- common formatting/context issues.
Force explicit records at each mode boundary
The report proposes a clearer separation format:- record an initial group version,
- record an AI-revised version,
- record a final human decision with justification.
Reconsider prompt templates
Prompt examples help reduce the initial barrier, but they can constrain thinking or embed assumptions. The authors suggest offering prompts at different “rigor levels,” requiring students to justify which one they used and how they adapted it.Plan infrastructure carefully
Anticipate limitations like access limits, file processing, context window constraints, and unavailability. Decide upfront whether personal accounts are allowed and how those interactions would be recorded (if at all).Add reflection checkpoints after each stage
After planning, search, screening, and extraction, groups should answer questions like:- What did we do without AI?
- What did AI suggest?
- What did we accept or reject?
- What remained uncertain?
This last piece is critical: it transforms AI usage from “faster production” into “methodological learning.”
A note about what the study can’t claim
Because this is an experience report, it doesn’t provide causal proof that LLMs improve SLR quality. It provides something arguably more actionable for teaching design: observed learning dynamics and risk patterns. For causal claims, the authors explicitly say they plan a controlled study comparing manual and LLM-assisted execution across matched groups.
Key Takeaways
- LLMs can reduce the initial barrier to producing SLR components like research questions, search strings, and criteria—especially in planning and search.
- AI also accelerates iteration, generating multiple alternatives quickly, which can make the classroom more exploratory.
- The biggest pedagogical risk observed isn’t just “wrong answers”—it’s over-delegation of structural decisions (questions, strings, inclusion/exclusion criteria, extraction structures).
- Students may show insufficiently critical acceptance, moving stage-to-stage after receiving plausible AI guidance without fully justifying methodological implications.
- Operational constraints matter: file/spreadsheet handling and access issues can cause students to spend time wrestling with tooling rather than practicing evidence synthesis.
- To make AI support real learning, instructors should add human supervision, decision records, structured stage checkpoints, and reflection moments—and avoid compressing the full SLR workflow into a single intense day.
- The study’s broader message for the future: AI should support the process, not conduct it. If you want auditable systematic methods, teach students how to supervise AI as part of their methodology—not as a shortcut around it.
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Observing the Conduct of Systematic Reviews with Generative AI Support: An Experience Report from a Graduate Software Engineering Course — arXiv
- Authors: Authors: Danilo Monteiro Ribeiro, Gilberto Sussumu Hida