Critical Thinking Gains Without Critical Thinking? Testing GenAI as a “Thinking Partner” in Stats

What happens when you use ChatGPT as a “thinking partner” in an undergraduate stats methods course? A constraint-first pilot found gains in statistical learning and AI literacy, yet critical thinking scores didn’t improve—because many students used the LLM like an answer checker.
The finding Students gained statistical learning and AI literacy, but their standardized critical thinking scores did not move in the pilot.
The method The intervention used a constraint-first thinking-partner protocol: think first, engage the AI, then report what changed.
The caveat Many learners still treated the LLM like an answer checker, so critical thinking gains depended on how the course positioned and assessed AI use.
1st MONTH FREE Basic or Pro • code FREE
Claim Offer

The Short Answer

In the constraint-first “thinking partner” pilot, students showed gains in statistical learning and AI literacy, but standardized critical thinking scores did not improve. The chat logs suggest many students used ChatGPT in a validator/answer-checker mode rather than pushing deeper reasoning.

Practically, this means course design and grading must target the reasoning process students do with the AI. If you only optimize correctness feedback, learners may become faster verifiers without strengthening transferable critical thinking.

A key caveat is that the pilot reported critical thinking outcomes staying flat even with improved learning, so “more use” or “better literacy” doesn’t guarantee higher-order reasoning transfer. Assess both learning and reasoning behaviors explicitly.

Critical Thinking Gains Without Critical Thinking? Testing GenAI as a “Thinking Partner” in Stats

Generative AI is already showing up in college classrooms, but there’s a real worry hiding underneath the excitement: will students use it in a way that replaces their critical thinking instead of strengthening it? A new study from the original paper takes a careful step toward answering that question by experimenting with ChatGPT as a “thinking partner” in an undergraduate research methods and statistics course.

What makes this research worth your attention is that it doesn’t just ask “did students like it?” The team ran a Design-Based Research pilot in Spring 2025 with N = 14 students, measuring learning, AI literacy, and critical thinking before and after the intervention—then digging into the student–AI chat logs to see how the tool was actually being used.

Why This Matters: The Real Risk Isn’t AI—it’s How We Let Students Use It

Here’s my take: the GenAI problem in education isn’t mainly “the model is too smart” or “students are lazy.” It’s more specific: without strong constraints, students default to using AI like an answer checker. And an answer checker can be helpful for correctness while still quietly shrinking the amount of thinking students do.

That’s exactly what this pilot suggests. Students showed big gains in statistical learning and AI literacy—yet their standardized critical thinking scores didn’t move. The logs help explain why: most students interacted with the LLM in a validator mode, where the model confirms or tweaks what they already wrote, rather than pushing them into genuine dialogue and reasoning.

A scenario you could apply today: imagine you teach a course with writing-heavy reflections (research methods, psych, sociology, even business analytics). You allow students to use a chat tool, but you don’t grade the quality of their reasoning process with the AI. What happens next is predictable: many students paste a draft question and ask, “Is this right?” The class gets faster feedback, but the critical reasoning muscle doesn’t necessarily get trained.

This research builds on a growing body of work showing that GenAI can reshape critical thinking into something like verification and integration—not the higher-order “challenge the assumption” kind that transfers to new contexts. The pilot gives that idea a concrete classroom mechanism: tool “positioning” and curriculum design decide whether students treat the model as a thinking partner or a rubber stamp.

What the Researchers Actually Built: A Constraint-First “Thinking Partner” Protocol in Stats

The intervention in the paper is designed around one core move: ChatGPT wasn’t positioned as an answer generator. It was framed as a thinking partner—closer to a dialogue partner than an authoritative source.

The three-step interaction protocol (and what “constraint-first” means)

Students followed a recurring interaction pattern:

  1. Think first (generate their own reasoning)
  2. Engage the AI (use ChatGPT to stress-test or refine their thinking)
  3. Report what shifted (describe how their thinking changed)

The “constraint-first” part matters because the designers didn’t rely on students figuring it out themselves. They introduced the protocol early and embedded it across repeated weekly activities, both in homework and in in-class practicums.

How the course and grading pushed the behavior

This wasn’t just “use ChatGPT if you want.” The course ran two parallel channels:

  • Homework assignments: students made methodological decisions using the thinking partner.
  • Weekly practicum sessions: students worked through problem sets alongside the LLM, with interaction documentation treated as routine.

Then came a crucial design detail: every assignment rubric had an AI partnership category worth ~5 points on average, and the rubric credited the quality of the student’s use of the thinking partner, not how impressive the AI output sounded.

So the course tried to make the “right use” of GenAI visible and gradeable—which, as we’ll see, had strong effects on statistical learning and AI literacy.

Two design revisions happened during the semester

DBR (Design-Based Research) isn’t one-and-done; it changes in response to what students actually do. The team made two mid-course revisions:

  1. Pedagogical revision: instead of expecting students to adopt the co-thinker stance on their own, practicum sessions rehearsed the same problem-solving rhythm that appeared in upcoming assignments.
  2. Epistemic revision: instructors reinforced that student decision-making mattered more than the LLM output—positioning the model as one input, not the destination of reasoning.

That matters because the qualitative findings later showed that many students were still using the model as a validator.

What They Measured: Domain Learning, AI Literacy, and Standardized Critical Thinking

The study used a pre–post design and three validated instruments—plus chat logs and instructor/student artifacts.

The measurement mix (and why it’s a big deal)

To answer whether the intervention affects thinking, you need more than one outcome lens. They used:

  • Statistical learning: AASCDM (10 self-report dimensions, including items about GenAI use)
  • AI literacy: MAILS (Meta AI Literacy Scale, 9 dimensions, self-report)
  • Critical thinking: WGCTA (Watson–Glaser Critical Thinking Appraisal, standardized), with subscales:
    • Drawing Conclusions
    • Evaluating Arguments
    • Recognizing Assumptions

Because WGCTA is designed to measure broader critical thinking, while AASCDM and MAILS measure course-aligned competence and self-reported AI literacy, the results can tell you whether gains transfer—or whether improvements are mostly localized to the course context.

Results summary: what improved and what didn’t

Here’s the key pattern:

  • AASCDM (statistical learning): students improved across all 10 dimensions.
  • MAILS (AI literacy): students improved on 8 of 9 dimensions.
  • WGCTA (critical thinking): no statistically significant change on any of the three subscales.

That combination—strong domain and AI-literacy gains, flat standardized critical thinking—isn’t a trivial result. It basically forces you to ask: what kind of “thinking” did students do with the AI, and what kind of thinking does the critical thinking test actually detect?

AASCDM: Gains across the board

For example, confidence in selecting statistical tests jumped from M = 1.57 to M = 2.79 (t = -6.50, p < .001). Students also reported large gains in:

  • using statistical software (M = 1.07 to 2.36, t = -6.62, p < .001)
  • explaining probability distributions (M = 1.64 to 2.93, t = -7.87, p < .001)
  • using AI tools (M = 1.36 to 2.86, t = -10.82, p < .001)

MAILS: AI literacy rose in most areas—except persuasion literacy

AI literacy improved on eight dimensions, including:

  • Apply AI: 36.64 to 53.00 (t = 5.82, p < .001)
  • Understand AI: 33.93 to 51.64 (t = 6.02, p < .001)
  • AI Ethics: 20.29 to 25.93 (t = 3.33, p < .01)
  • AI Emotion Regulation: 23.64 to 26.50 (t = 2.46, p < .05)

But one dimension didn’t move:
- AI Persuasion Literacy: 25.43 to 25.86 (t = 0.32, p > .05)

That “flatness” is important because persuasion literacy often requires exactly the kind of evaluative, transfer-heavy reasoning that standardized critical thinking measures try to capture.

WGCTA: Critical thinking scores didn’t budge

On the WGCTA subscales, none of the pre–post changes were significant (all p > .05). For instance:

  • Drawing Conclusions: 44.43 to 36.71 (change -7.71, p = .199)
  • Evaluate Arguments: 36.14 to 35.86 (p = .987)
  • Recognize Assumptions: 32.71 to 36.50 (p = .451)

So the intervention increased competence and AI literacy, but did not measurably raise standardized critical thinking within a single semester.

Now for the most revealing part: the team didn’t just measure outcomes—they analyzed 442 student-authored prompts across seven assignments. The chat logs included a mix of questions about course methods (z-scores, hypothesis tests, sampling bias) and meta-level prompting (asking for feedback, reflection, and critique).

What kinds of prompts students sent

They found nine topic clusters, including statistical methods and research design concepts, plus discourse topics like feedback and reflection.

More importantly, their coding categorized prompts by function. Two broad groups dominated:

  • Context Description (background, acknowledgment): 44.1% of messages
  • Information Requests/Needs (requests for feedback/help/reflection/higher-order thought): 55.9%

Students also tended to “stay in a mode” rather than switching interaction styles mid-conversation. The biggest prompt-to-prompt transitions were repetitive loops like:

  • assignment description → assignment description (21.9% of turn transitions)
  • general improvement request → general improvement request (10.1%)

That suggests a learning pattern more focused on continuity of task completion than iterative conceptual challenge.

The three “positionings” that explain the critical thinking result

Qualitative analysis produced one central framework: how students positioned the LLM.

  1. Answer generator: students offload the thinking (“Here’s the reflection question—answer it.”)
  2. Answer validator (most common): students draft answers and ask the AI to confirm or check them.
  3. Co-thinker: students treat the model like a partner in reasoning, asking follow-ups, stress-testing assumptions, and building the question with the AI.

And here’s the critical finding: the modal relationship was validator, not co-thinker.

What interaction looked like in each mode

  • Validator mode often looked like: paste draft → “Are my answers correct?” → copy-paste AI response with little visible integration.
  • Generator mode looked like: minimal prompting with the expectation the AI will generate the answer.
  • Co-thinker mode appeared in fewer interactions and showed real multi-turn reasoning—students sharpening questions and expanding their thinking.

A subtle design effect: AI responses encouraged “closure”

The researchers also noticed something small but potentially powerful: the model often ended responses with closed questions like “Does that make sense?” or “Would you like me to explain further?” Students often replied “yes” and moved on.

That interaction style can unintentionally steer students back toward validator behavior: confirm, close, proceed.

The team even considered adjusting system prompting to elicit more open-ended follow-ups to keep dialogue—and critique—alive.

Connecting it back to the results

This is the bridge between the data:

  • Students improved statistical learning and AI literacy because the course scaffolding and repeated use strengthened domain skills and tool familiarity.
  • But standardized critical thinking didn’t change because most students weren’t consistently practicing the challenging, dialogic reasoning that standardized tests tend to detect.
  • The persistence of validator-style use is the mechanism that likely limited transfer.

The paper’s own synthesis states this plainly: gains happened where tasks were scoped and supported; critical thinking gains lagged where students defaulted to validation instead of dialogue.

What to Do Next: Design Principles for Actually Training Critical AI Literacy

The authors finish with six design principles generated from their DBR iteration. Here are the ones that feel most actionable if you’re teaching (or designing) GenAI-enabled courses.

1) Treat the system prompt as instructional design, not a generic setting

The system prompt can’t be improvised at the start. It must fit the course and assignments ahead of time.

2) Reflection can’t be an occasional add-on

If reflection shows up only in one or two tasks, students won’t reliably develop co-thinker engagement. The next iteration integrated reflection more broadly across assignments.

3) Rubrics must grade the thinking process with the AI, not just outputs

Since the pilot rubric already had an AI category, the next step is making co-thinker behavior more clearly rewarded—especially visible critique and extension.

4) Instructor prep is required, not optional

Constraint-first design needs assignment redesign and rubric calibration before the course begins.

5) Weekly instructor support and adjustment matters

This pilot had one instructor iterating, but scaling likely requires shared spaces to interpret student reflections and tune instruction quickly.

6) Student reflections are design data

Instructors should treat reflections diagnostically: evidence for what’s working and what’s failing in the design, not just something to grade.

A practical “today” takeaway

If you want critical thinking growth in GenAI-supported courses, don’t just require students to use the tool. Require them to disagree with it in visible ways—and make that disagreement part of the grade. Validator behavior is comfortable and fast; co-thinker behavior is harder, slower, and requires structured incentives.

And importantly: this study also suggests a specific gap—AI persuasion literacy didn’t improve, and that overlaps strongly with the kind of higher-order critical thinking the standardized assessment targets.

Key Takeaways

  • In this pilot (N = 14, Spring 2025), students’ statistical learning improved across all 10 AASCDM dimensions, and AI literacy improved on 8 of 9 MAILS dimensions.
  • Standardized critical thinking (WGCTA) did not show measurable gains on any of its three subscales.
  • The qualitative “why” is strong: most students positioned the LLM as an answer validator rather than a co-thinker.
  • Interaction patterns mattered: students often used the AI to confirm drafts (“Is this correct?”) and then moved on, which reduced opportunities for challenging reasoning.
  • AI persuasion literacy was the one MAILS dimension that didn’t change, aligning with the idea that evaluative, transfer-heavy thinking needs more explicit instruction and practice.
  • Practical implication: if you want critical thinking—not just correctness—design rubrics and assignments that reward visible critique, disagreement, and dialogic reasoning with the AI, not just checking answers.

If you want, I can also turn these design principles into a ready-to-use checklist for a syllabus/assignment rubric (homework + reflection + AI log requirements) based on the structure used in the paper.

Sources Used

This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:

Where To Go Next

Generative AI Isn’t Ready to Replace Stats Experts—But It Can Help

Title: Multi-Agent Defense: Securing AI Pipelines Against Modern Threats

Can AI Really Code Stats? Testing ChatGPT & Llama on SAS Programming

Browse the free Prompt Database or tune your own prompts with the Prompt Optimizer.

Frequently Asked Questions

Limited Time Offer

Unlock the full power of AI.

Ship better work in less time. No limits, no ads, no roadblocks.

1ST MONTH FREE Basic or Pro Plan
Code: FREE
Full AI Labs access
Unlimited Prompt Builder*
500+ Writing Assistant uses
Unlimited Humanizer
Unlimited private folders
Priority support & early releases
Cancel anytime • 10,000+ members
*Fair usage applies on unlimited features to prevent abuse.