sözaltı news Science
Science
EN AZ
Research shows how AI is getting better at exams: How universities can respond

Research shows how AI is getting better at exams: How universities can respond

phys.org 30.09.2026 16:00 2 views
When ChatGPT arrived at the end of 2022, universities scrambled to determine whether generative AI was good enough to pass assessments and make it easy for students to cheat.

This article has been reviewed according to Science X's editorial process and policies. Editors have highlighted the following attributes while ensuring the content's credibility: When ChatGPT arrived at the end of 2022, universities scrambled to determine whether generative AI was good enough to pass assessments and make it easy for students to cheat. There were some big claims by AI companies at the time, including that tools showed "human-level performance on various professional and academic benchmarks." Headlines boasted about AI passing U.S. bar exams.

But early evidence suggested generative AI did not always live up to the hype. Our previous research showed models performed poorly on tasks requiring the essential skill of critical legal analysis. In other words, the AI models would struggle to compete with a real law student in legal analysis.

But as AI models get more sophisticated, we wanted to know whether that was still true. Our 2023 experiment tested generative AI on a real Australian criminal law exam. It found "generative AI isn't close to replacing humans in intellectually demanding tasks such as [an undergraduate] law exam." Our new study revisited that experiment.

This time, we tested nine models from five AI providers across two compulsory law subjects at the University of Wollongong. Most of the AI tools are easily accessible via subscription. Once the students' criminal law and tort law (which deals with compensation) final exam questions had been prepared, we used each model to generate an answer to the same exam papers, producing 18 answers in total.

The models received the exam questions and prompts, but no subject-specific lecture notes, textbooks or curated legal materials. They did have internet access, and some had an "enhanced reasoning" feature. This gives the AI more time to plan, consider different approaches and check its reasoning before answering.

Six of the 18 papers were graded by subject coordinators (the authors here), and the remaining 12 were mixed among genuine student papers and blind-graded by tutors who did not know anything about the involvement of AI models. Compared with 2023, the improvement was significant. In criminal law, AI papers averaged 76.3%, outperforming 82.5% of students.

Extract — continue reading at the source.

Read full story