The Short Answer
Course context appears to shape how learners use GenAI in programming MOOCs more than learner demographics do, affecting both how often and what learners use it for (like debugging or code explanation).
Practically, this means MOOC guidance can be built around the course’s real activities and common learning challenges, rather than targeting advice by age, gender, education, or prior experience.
Because the study relies on post-course, self-reported questionnaire data from two specific Estonian Python MOOCs, results may not generalize to every platform or programming language.
On this page
- Introduction: What happens when ChatGPT meets real MOOC learners?
- Why this matters: Course design is becoming the real “AI policy”
- What the researchers actually measured: adoption, frequency, and purpose (not just “used GenAI”)
- Adoption was common, but usage frequency and purposes shifted with course length
- Debugging and idea generation increased in the longer course—purpose matters more than “AI use exists”
- Demographics didn’t predict GenAI use—so what does predict it?
- GenAI use didn’t clearly translate into better (or worse) measured outcomes
- What course designers should do next: build support around context, not stereotypes
- Key Takeaways
GenAI in Programming MOOCs: How Course Context Shapes Use
Introduction: What happens when ChatGPT meets real MOOC learners?
If you’ve ever watched students struggle with a coding assignment, you already know the pattern: they want answers fast, but they also want to understand why the answer works. Now add generative AI tools—think ChatGPT and code-completion assistants—and that old pattern changes. A big open question is how learners actually use generative AI inside programming MOOCs, and whether it mostly depends on who the learner is (age, experience, education) or on what the course is like.
New research from Marina Lepp explores exactly that. The study (based on the original paper) looks at two Estonian-language programming MOOCs that both teach Python, but differ a lot in length, workload, and how many assignments students do. The key idea: GenAI use might be widespread, but the “shape” of that use could depend more on the course context than on learner demographics.
The study uses post-course questionnaire data from 187 participants in “About Programming” and 182 participants in “Introduction to Programming,” measuring (1) whether learners used GenAI at least once, (2) how frequently they used it, and (3) what they used it for—like debugging, idea generation, code explanation, code solutions, or answering tests. Let’s unpack what the research found and what it means for anyone designing, teaching, or learning programming online.
Why this matters: Course design is becoming the real “AI policy”
This research is significant right now because MOOCs are no longer blank slates. They’re living ecosystems where learners already have GenAI available on their phones and laptops—whether instructors plan for it or not. And since this study found minimal differences by gender, age, education level, or prior programming experience, the usual instinct—“just target the guidance to certain demographics”—doesn’t look like the most powerful lever.
Here’s a practical scenario you could apply today: imagine you’re revising a programming MOOC for beginners. You might be tempted to write one blanket rule like “Don’t use AI to cheat.” But this paper suggests a more actionable approach: build guidance around common GenAI purposes that map onto your course’s real challenges, especially debugging and idea generation. In other words, the course itself is shaping GenAI behavior—so your support should shape it too.
This also builds on earlier AI-in-education research that often shows mixed learning outcomes. A lot of studies zoom in on whether students used AI at all, but this paper makes a sharper move: it distinguishes adoption vs. frequency vs. purpose, and then shows that frequency/purpose differ by course context. That’s a methodological upgrade—and it fits the broader trend in AI education research that emphasizes how learners interact with tools, not just whether they do.
What the researchers actually measured: adoption, frequency, and purpose (not just “used GenAI”)
Lepp’s study focuses on self-reported GenAI use during two MOOCs:
- MOOC1: “About Programming” — 4 weeks, expected workload 26 hours, n = 187 respondents
- MOOC2: “Introduction to Programming” — 8 weeks, expected workload 78 hours, n = 182 respondents
Both courses were taught in Estonian and used Python.
The study tracked multiple layers of behavior:
The study’s GenAI metrics
- Adoption (Yes/No): Did the learner use GenAI at least once?
- Overall usage frequency: How often, using categories from “once” to “weekly”
(Non-users were treated as frequency0when computing overall comparisons.) - Purpose: Among learners who used GenAI, what did they use it for?
- debugging
- idea generation
- code explanation
- generating code solutions
- answering weekly tests
- Purpose-specific frequency: For each purpose selected, how often they used GenAI for that purpose?
The learning outcomes they checked
They also pulled performance and activity indicators from the learning management system, including:
- mean weekly test score (max 10 points per test)
- mean number of submission attempts per assignment
- numbers of completed choice exercises and optional exercises
Notably, all assignments and tests were unproctored (done at home), so results reflect performance under course conditions.
The big comparison challenge
A crucial detail: MOOC1 and MOOC2 differ in duration, workload, topic complexity, and assignment volume all at once. So the results strongly suggest “course context matters,” but they can’t isolate which single factor (like longer duration vs. more assignments) caused the differences.
Still, that’s often how course design really works—everything changes together—so the findings are practically useful.
Adoption was common, but usage frequency and purposes shifted with course length
Let’s look at the headline findings first: GenAI adoption was widespread in both MOOCs. But the frequency and how learners used GenAI were more course-dependent.
Adoption (used at least once)
- MOOC1: 142 / 187 = 75.9%
- MOOC2: 150 / 182 = 82.4%
The difference was not statistically significant (so adoption was basically “high everywhere”).
Usage frequency (how often they used GenAI)
Even though adoption was similar, frequency wasn’t.
| Metric | MOOC1 (“About Programming”) | MOOC2 (“Introduction to Programming”) | What changed |
|---|---|---|---|
| Median overall frequency | 2 | 2 | Same median, but distribution differs |
| Mean overall frequency | 1.57 (SD 1.27) | 2.08 (SD 1.40) | MOOC2 higher |
| Weekly/near-weekly users | 10.2% | 18.7% | More frequent in MOOC2 |
| Stat sig? | — | Yes (p < 0.001) | Higher frequency in longer course |
The effect remained significant even when looking only at learners who reported using GenAI (so it wasn’t just “more adopters” in MOOC2—it was more frequent usage among users).
Which tools students mentioned
ChatGPT dominated in both courses:
- MOOC1: 88.7%
- MOOC2: 82.0%
Other tools appeared more in MOOC2 (for example Bing/M365 Copilot: 14.7% vs 7.0% in MOOC1). Less-common tools were mentioned only by small percentages.
Practical implication
If you’re designing support for GenAI in MOOCs, you can’t assume “same adoption pattern means same usage pattern.” Even with identical general guidance rules across courses, learners may still adjust their GenAI behavior depending on workload and assignment demands.
Debugging and idea generation increased in the longer course—purpose matters more than “AI use exists”
This is where the study gets really interesting. It’s not enough that learners used GenAI—it matters what they used it for, because those purposes likely map to different learning behaviors.
Purpose prevalence differed between MOOCs
Among GenAI users, MOOC2 showed higher prevalence for two purposes:
| Purpose | MOOC1 n (%) | MOOC2 n (%) | Stat result |
|---|---|---|---|
| Debugging | 108 (76.1%) | 133 (88.7%) | p = 0.023 |
| Idea generation | 38 (26.8%) | 63 (42.0%) | p = 0.040 |
| Code explanation | 69 (48.6%) | 65 (43.3%) | not significant (p = 1.000 after correction) |
| Generating code solutions | 16 (11.3%) | 18 (12.0%) | not significant |
| Answering weekly tests | 4 (2.8%) | 8 (5.3%) | not significant |
So learners in the longer course were more likely to report GenAI for debugging and brainstorming/idea generation.
Purpose-specific frequency also shifted (sometimes differently than prevalence)
Now look at how often they used GenAI for specific purposes. In MOOC2, frequency was higher for some purposes:
| Purpose | MOOC1 Mean freq | MOOC2 Mean freq | Stat result |
|---|---|---|---|
| Debugging | 2.10 (0.98) | 2.64 (1.12) | p < 0.001 |
| Code explanation | 2.30 (1.08) | 3.08 (1.16) | p < 0.001 |
| Idea generation | 2.11 (1.13) | 2.37 (1.13) | not significant after correction |
| Code solutions | 2.25 (1.29) | 3.28 (1.02) | not significant after correction |
| Answering tests | 2.25 (1.89) | 3.13 (1.36) | not significant after correction |
A key nuance the paper emphasizes: frequency and prevalence don’t always move together. For example, code explanation didn’t show a prevalence difference, but users in MOOC2 reported higher frequency for explanation. So if someone only asked “Did you use GenAI for code explanation?”, you’d miss how usage intensity differs.
Practical implication
For instructors, “GenAI policy” should be more like “GenAI coaching.” If debugging is much more common in your course, build debugging-oriented guidance: how to verify AI output, how to write tests, how to compare alternatives, how to reflect on why a fix works.
Lepp’s work also naturally connects back to the larger discussion in the original paper: purpose and frequency are distinct, and course context can shift both.
Demographics didn’t predict GenAI use—so what does predict it?
One of the most surprising (and useful) findings: within each course, adoption and frequency didn’t significantly differ across:
- gender
- age group
- education level
- prior programming experience
That doesn’t mean the relationships don’t exist in every setting—statistically, they just didn’t show up here. The paper also notes effect sizes were generally small, and the sample is heterogeneous (mean age around 37.7, majority female at about 71–73%).
What this suggests (carefully)
Because access to GenAI is easy and broadly available, it may reduce some barriers that usually create demographic gaps. But the study didn’t measure things like “perceived ease of use,” so we can’t confirm that mechanism.
Practical implication: stop over-personalizing the rule
If you’re creating course-level guidance like “older learners shouldn’t use GenAI” or “novices are more likely to rely on it,” this paper argues those demographic-targeted assumptions may be off-base. Instead, focus on task-level and course-level structure, because context seems to be doing a lot of the shaping.
GenAI use didn’t clearly translate into better (or worse) measured outcomes
This study didn’t find strong evidence that GenAI use improves test performance or engagement—at least not in the metrics it examined.
Users vs non-users: mostly similar performance
In MOOC1, there was one statistically significant difference:
- non-users completed more optional exercises than GenAI users
(mean 2.09 vs 0.96; correlation reported as r = −0.251, adjusted p = 0.027)
But test scores and submission attempts didn’t show significant differences.
In MOOC2, none of the measured outcomes differed significantly between users and non-users after correction.
Frequency vs performance: no significant correlations
They also checked whether more frequent GenAI use correlated with better or worse performance. In both courses, correlations ranged from about −0.139 to 0.117, and none were statistically significant after correction.
Interpreting this without jumping to conclusions
There are a few plausible explanations:
- Learners may use GenAI in ways that don’t map cleanly to these performance indicators.
- GenAI might be helping with confidence, explanation quality, or learning progress that doesn’t show up in unproctored weekly tests.
- Or GenAI use might be compensating for learning difficulties, leading to “similar outcomes” rather than improvement.
The study’s careful language is important: frequency and adoption don’t guarantee effective learning support. Prior work (and common sense) suggests quality depends on the interaction style—like whether learners ask for debugging steps and then verify.
The authors also point out a limitation: this study used self-report, not direct logs of how GenAI was used.
What course designers should do next: build support around context, not stereotypes
If you’re responsible for a programming MOOC (or planning one), here are concrete moves suggested by this research’s pattern: context shapes usage, and usage doesn’t automatically improve measured outcomes.
1) Treat debugging as a first-class learning activity
Since MOOC2 saw higher debugging prevalence and frequency, design assignments so learners must demonstrate reasoning, not just produce a fix. For example:
- require tests or minimal reproducible examples
- ask learners to explain what caused the bug and how they validated the fix
- include “debugging reflection” prompts (short, but frequent)
2) Separate “using AI” from “understanding AI output”
Because purpose doesn’t guarantee depth, your guidance should force verification behaviors:
- check for edge cases
- compare AI output with at least one alternative approach
- require “what I tried first” summaries (even in a forum)
3) Don’t write one-size-fits-all rules—write course-specific guidance
This paper suggests that different course structures lead to different GenAI use patterns. So if your course is longer or has more assignments, you should expect more frequent GenAI use and adjust scaffolding accordingly.
4) Collect better evidence than questionnaires when possible
Future research (and course evaluations) should combine self-report with interaction data—like logs or structured prompts—to see whether GenAI is used for iterative refinement vs. copy-paste solutions.
Key Takeaways
- GenAI adoption was high in both MOOCs: 75.9% (MOOC1) vs 82.4% (MOOC2), but the difference wasn’t statistically significant.
- Course context mattered for usage intensity: learners in the longer, more extensive MOOC (MOOC2) used GenAI more frequently (including more weekly/near-weekly use: 18.7% vs 10.2%).
- Purpose shifted with course length: MOOC2 users were more likely to use GenAI for debugging and idea generation, and—among users—also showed higher frequency for debugging and code explanation.
- Learner demographics didn’t predict GenAI use here: no significant differences by gender, age, education level, or prior programming experience within each course.
- Measured learning outcomes didn’t strongly improve with GenAI use: in most comparisons, users and non-users looked similar on tests/submissions; frequency wasn’t significantly correlated with performance after correction.
- For instructors: focus guidance on common course realities (especially debugging) and on verification and reflection, not just broad “don’t use AI” rules.
- For future MOOC research: don’t rely only on “used GenAI yes/no.” Track purpose and frequency, and ideally add behavioral logs to understand the learning impact.
If you want, I can also turn these findings into a practical “GenAI-ready MOOC checklist” for instructors (what to change in assignments, rubrics, and student-facing guidance).
Sources Used
This article is a plain-English breakdown of the following peer-reviewed preprint. Read the original for full methodology and results:
- Understanding Generative AI Use in Programming MOOCs: The Role of Course Context and Learner Characteristics — arXiv
- Authors: Authors: Marina Lepp