Post

LLMs Pass Machine Learning Finals and Generate New Exam Questions

We collected machine learning final exams from MIT, Harvard, and Cornell and threw GPT-3, OPT, Codex, and ChatGPT at them. These weren’t just single answer problem sets. They’re multi-part, multi-topic finals that take professors days to write. We found that translating the questions into code and providing few shot examples dramatically improved the problem solving ability of the models.

More importantly, we had the models generate new exam questions and ran a student survey to see if they could tell them apart from professor-written ones. Across quality, difficulty, and appropriateness, the AI-generated questions were indistinguishable from the human generated questions.

The part we cared about most was the implication for teaching. Banning ChatGPT from classrooms seemed like the wrong move. Students could use these models to get a custom curriculum, generating practice questions tailored to the specific topics they’re struggling on rather than just memorizing a fixed set of problems. We hope this work helps push AI’s influence on education in a positive direction.

Read the full paper: Generating and Solving Machine Learning Finals with Large Language Models

This post is licensed under CC BY 4.0 by the author.