What kind of AI co-working leaves students understanding more without AI?
Dinara Pisareva · Nazarbayev University · Principal Investigator
Denis de Crombrugghe · Narxoz Business School · Co-Investigator
Sustained intellectual work with AI as a genuine partner in a shared system of thinking.
Teaching cannot act on the student–AI configuration itself, but it can develop what the student brings to it, so we measure the student-side orientation — three co-equal dimensions:
Working with AI to think further, not to finish faster.
Bringing curiosity and imagination — trying things you hadn't thought to try.
Pushing back and building over many rounds, not taking the first answer.
Ontological openness — staying open about what AI is — encourages all three: taught and measured alongside them, but not a facet of the orientation.
"My goal with AI is to think further than I could alone."
"I work with AI to come up with options I wouldn't have thought of alone."
"My best work with AI happens over several rounds, not one."
"I treat what AI is as an open question."
A within-person pre/post design embedded in three authentic courses at Nazarbayev University, all taught with the same partnership pedagogy:
| Course | Who | AI work students do |
|---|---|---|
| Research Methods (PLS210) | Undergraduates | Weekly AI conversations; research designs co-built with AI; creative presentation of research projects on GitHub. |
| AI & Social Science (PLS419/519) | Upper-level + MA | A semester-long public project built with AI and presented on GitHub, plus two op-eds about AI. |
| Qualitative Methods (PLS514) | Graduate students | Analytical portfolios: grounded theory, RTA, and process tracing applied to AI-generated case studies, first without and then with AI, plus a comprehensive comparison of the two passes — differences, commonalities, implications for data interpretation. |
All measures run in Weeks 1 and 14; every model adjusts for the student's own baseline.
Twelve self-developed CP items (three per dimension plus the antecedent block), the AIDep-22 Cognitive Dependence subscale (Wu et al., 2026), and an adapted 25-item AI-literacy scale (Lee & Park, 2024).
Sixteen course-specific items per wave (12 multiple-choice, 4 open-ended) targeting understanding and application; half repeat across the two waves, half are wave-specific.
The quiz is the study's objective outcome — the residue measure.
The predictor is the partnership composite — the mean of purpose, engagement, and iteration; the three tests are Holm-corrected. H1 carries the evidential weight — its outcome is objective; H2 and H3 are self-report and triangulate.
Controlling for Week-1 levels, the partnership composite positively predicts no-AI knowledge-test scores at Week 14.
Controlling for Week-1 levels, the partnership composite negatively predicts cognitive dependence on AI at Week 14.
Controlling for Week-1 levels, the partnership composite positively predicts total AI-literacy scores at Week 14.
Powered at .80 for medium effects (r ≈ .33; .38 at the strictest Holm bar); smaller effects arrive as wide intervals, reported as such. Per-dimension effects and the antecedent→composite link are exploratory.
Sixteen course-specific items per wave (6 + 2 repeated across waves): 12 multiple-choice (0/1) + 4 open-ended scored 0–2 against a rubric by two independent raters; total 0–20, z-scored within course.
The Cognitive Dependence subscale of the AIDep-22 (Wu et al., 2026), five items administered intact.
An adapted 25-item AI-literacy scale (Lee & Park, 2024): the total score is confirmatory; the five subscales are an exploratory breakdown.
Analysis: regression-adjusted change (ANCOVA) with course fixed effects and prior AI use as covariate, robust standard errors; estimation-led reporting with 95% CIs.
The design must answer its own strongest objection: the researcher built the framework, teaches it, and measures it. The answer is architectural.
Hypotheses, instruments, and the full analysis plan are locked on OSF before the first class meets.
A co-investigator holds all survey responses and consent forms during the semester; the PI does not know who consented into the study until final grades and appeals close.
Two raters score independently against rubrics written before any data; agreement statistics and an instructor-only sensitivity analysis are reported.
Participation is voluntary and pseudonymous; consent decisions are invisible to the instructor and cannot touch grades. Consent self-selection likely tilts the sample toward engaged students — range restriction on the predictor, which makes the tests conservative.