The Generative AI Learning Penalty: Evidence from Chinese Secondary Education

Using 30 months of panel data on 26,811 Chinese students in grades 7-12, we study how generative AI affects homework productivity and learning. The data combine monthly closed-book exams, high-school and college entrance exams, and homework scores and completion time across nine subjects. We exploit staggered AI adoption in a difference-in-differences design. AI adoption raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months. High-stakes entrance-exam scores fall by 18 and 24%, with the full penalty emerging only after about two years. The losses are largest in social science subjects, followed by STEM and languages, and are especially large for junior students, high-achieving students, and boys. The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores. AI users who maintain similar homework completion time as non-AI users experience small learning losses.

  • SootySootySoot [any]@hexbear.net
    link
    fedilink
    English
    arrow-up
    37
    ·
    edit-2
    4 days ago

    AI users who spend as much time on homework as non-users achieve similar exam scores, even though their higher homework scores indicate that they use AI. These students are not differentially selected on prior achievement, suggesting that generative AI does not reduce their learning efficiency.

    I think this is interesting. The study is (sort of) correlative so it’s not perfect. But their findings suggest that the big driver of the difference is simply that students spend less time on their homework, rather than insurmountable brainworms.

    That being said, it’s correlative. It could just be that students who spent less time on homework did, in general, offload more of their thinking to AI and also ingrain themselves with brainworms.

    Actually, doing a heel turn, they also say those who use AI impaired their future learning as well. Which does lend meaningful credence to the brainworm theory.

      • Soot [any]@hexbear.net
        link
        fedilink
        English
        arrow-up
        19
        ·
        edit-2
        4 days ago

        Yes, there is, and it’s practically an exact correlation regardless of AI use:

        Again though, very much worth nothing that this is correlative, there’s no RCT here in regards to AI use or homework time.

        See also some other (not overly well formatted) graphs showing relationships. Note interestingly that the “Pre-AI” and “Never-AI” groups are nearly identical:

        More direct link to the paper if it’s any use

        • Soot [any]@hexbear.net
          link
          fedilink
          English
          arrow-up
          19
          ·
          4 days ago

          Actually, on searching more graphs, there’s what I’d argue is a directly more interesting graph. That the hours a student spends on general AI use per week strongly correlates with declined exam scores.

    • DasRav [any, any]@hexbear.net
      link
      fedilink
      English
      arrow-up
      6
      ·
      4 days ago

      I contend that these students are idiots. The point of AI is to not to the work. If you use AI and then learn it all anyway, why are you using AI? have some self-respect.

  • EmmaGoldman [she/her, comrade/them]@hexbear.net
    link
    fedilink
    English
    arrow-up
    34
    ·
    4 days ago

    AI is causing the single largest degradation in human intelligence ever recorded.

    The black plague killed 20% of the global population and AI is having an even larger impact on our total collective intelligence than 1 in 5 people ceasing to exist.

    • Infamousblt [any]@hexbear.net
      link
      fedilink
      English
      arrow-up
      29
      ·
      edit-2
      4 days ago

      I see this in my day to day. I have one of those jobs where I’m being tracked on how much AI I use and I have to use at least some or I get nasty questions, so I use it anytime I need to do something mundane that I could do without it (this way it’s easy for me to check that it did it right too and fix the mistakes). The other day I went to use it for some mundane bullshit work thing but the server was down. So I opened up PowerPoint or whatever and then I was like wait, how do I do this? There was that brief moment where my brain had to go digging for the knowledge.

      That moment was terrifying to me. It’s shit I know how to do. I’ve done it for years. My whole career. And yet I went to do something I’ve done hundreds of times and for a split moment I forgot how, because I’ve been letting the lying machine do it for me so I can meet my dumbass AI quotas.

      Within like a minute I had fully pulled out all my “here’s how to PowerPoint” knowledge and it was fine and everyone was like wow great job or whatever who cares it’s bullshit mundane work things.

      But it really drove home that the only reason I was OK in that moment was because I’ve done this shit for years.

      There’s now two whole generations now that never learned how to do mundane work shit. The boomers are on their way out, but this young AI raised generation is also going to have no clue how to adjust the color on a slide deck because they are so used to the AI just doing it for them. And they’re going to be just like boomers when the AI is down or is taken away from them. They’ll have no clue what to do, but more importantly they’ll have no clue how to figure it out.

      If something isn’t done about this at a young age we might be about to live in a world where the vast majority of people never learn how to think. That’s the most dangerous thing of all.

    • hello_hello [hy/hym]@hexbear.net
      link
      fedilink
      English
      arrow-up
      13
      ·
      4 days ago

      Using AI is already a trip and a half when you’re an expert in whatever domain subject you’re referencing it with. Computer programmers are lucky enough to have tools for static verification.

      Using LLMs as a student sounds like a nightmare. I hope Chinese regulators crack down on LLMs before serious harm gets widespread.

    • SootySootySoot [any]@hexbear.net
      link
      fedilink
      English
      arrow-up
      25
      ·
      edit-2
      4 days ago

      Pure speculation on my part, but my understanding is that East-Asia as a region generally embraces a mindset of giving students problems that will frequently be too hard to solve, and students are simply expected to get as far as they get with it. I’m guessing that puts a theoretical maximum score cap way higher than even the smartest kids should ever realistically achieve.

      • Blockocheese [any]@hexbear.net
        link
        fedilink
        English
        arrow-up
        20
        ·
        4 days ago

        They started implementing something similar to this at my high school in the US like 2 years before I graduated

        For chemistry and biology they’d have labs where to get the correct answer, we’d need to utilize something we hadn’t been taught before and the teacher wasn’t going to tell us.

        When the lab was over they’d tell us the correct way to figure it out which was pretty discouraging to me because I hated being graded on something they knew I didnt know.

        The correct answer wasn’t a large part of the grade and they said students learn better when they fail like that at first but I fucking hated it angery

    • AstroStelar [he/him]@hexbear.net
      link
      fedilink
      English
      arrow-up
      3
      ·
      edit-2
      3 days ago

      100 was set as the average score among students that don’t use AI.

      I looked through the study and the graphs are very confusing to read, but the conclusion is very concise so I’ll post it here:

      Conclusion and Discussion

      Using large-scale evidence from a natural school setting, we find that generative AI use substantially reduces learning among secondary-school students. The central pattern is stark: AI raises homework scores and reduces homework time, but lowers performance on closed-book exams. Regular monthly exam scores fall by 20 percent, while high-stakes entrance-exam scores fall by 18 and 24 percent.

      The negative learning effects fully materialize after six months for regular exam scores and after two years for entrance exams. The negative effects on learning outcomes appear to be mostly driven by the 81 percent of AI-using students, who spend less time on homework than even the fastest non-AI student, receive high homework scores matching the capability of generative AI tools they are using, and yet very low exam scores. By contrast, AI students who spend as much time on homework as non-AI students achieve similar exam scores and higher homework scores. [However, almost no AI-using students spend this much time on homework anymore after more than five months of use.]

      Our findings highlight the importance of the demand side and student incentives. Much of the existing RCT literature focuses on the supply-side question of how to design AI tools that scaffold learning while limiting outsourcing. In China, as in many other countries, such tutoring tools already exist and are often available at low or zero cost. Yet most students in our setting do not choose them. Instead, they rely on general-purpose AI tools that provide quick answers and facilitate homework outsourcing.

      How can we incentivize students to choose tutoring AI tools in an environment with general-purpose AI tools that give direct answers? Our findings suggest several possible interventions. First, providing credible information about the long-run learning costs of homework outsourcing may alter student behavior. Second, increasing the weight placed on closed-book, in-person assessments may help restore the connection between effort and reward. Third, parents and teachers may be more effective if they monitor inputs, such as homework time and study effort, rather than outputs such as homework scores.

      I also found the section on who gets most affected interesting:

      • Effect is worse in junior high than high school, suggesting uninformed (over)use makes things worse;

      • “The average estimated effect 6-10 months after adoption for students who report using generative AI 0-1 hours per week is -5 percent, compared to -30 percent for students who report using generative AI 5 hours or more.”

      • The effect is slightly worse for boys than for girls, widening the gender gap (boys on average score worse).

      • High-scoring students who adopt AI see the largest declines, poor-performing students are affected less.