Frankhuang/iStock/Getty Images Plus
Articles and essays published over the past year have correctly observed that the wide availability of AI means we need to go back to proctored exams, where students must demonstrate their knowledge under controlled conditions. But we strongly disagree that going back to handwritten answers in blue books, as many propose, is the best path forward. Advances in technology make computer-based testing a far superior alternative, provided the exams have certain characteristics. AI makes the case for computer-based testing even stronger.
The term “computer-based testing” may evoke images of thousands of students simultaneously taking a standardized exam such as the MCAT, LSAT or GRE, filling in multiple-choice answer bubbles or simultaneously answering the same set of essay questions. But technology in use at a number of leading universities, including our own, enables computer-administered assessments that are richer and more authentic. Of course, for open-ended essay questions, typing an essay on a computer is no harder for most students than writing in a blue book, and may even be easier.
But beyond this, anything that can be displayed in a web browser—images, video clips, graphs, formulas—can be included in a question. Math problems can ask students to sketch parts of a graph using the mouse or by dragging and dropping graph segments on the screen. Physics problems can show arbitrary arrangements of pulleys, ropes and weights. Computer coding questions allow students to solve problems involving large and realistic datasets using the same coding tools professionals use. The possibilities are limited only by the capabilities of a web browser, which today allow for images, animated diagrams, video clips, drag-and-drop, drawing and much more.
With little effort, computer-based exams can generate unique versions for each student. An instructor can create a handful of variations of a question that test the same knowledge component in different ways. In many domains, especially STEM disciplines, short bits of computer code can be written to automatically generate many slightly different variations of a question. For example, multiple variants of a question about electrical circuits could feature different arrangements and values of the circuit components.
The availability of multiple question variants allows for the creation of an exam that is practically unique to each student, one that’s effectively equivalent in the knowledge being tested but largely neutralizing the benefit of hearing “what’s on the exam” from a fellow student.
Generating a fair and unique exam for each student means that students can self-schedule when to take an exam, at their convenience, over a multiday window. They do so in dedicated labs outfitted with institutionally managed computers and staffed by human proctors up to 80 or more hours per week. The proctors are specialized employees (often students on our campuses) who undergo training and have more experience supervising exams than most instructors or teaching assistants. Proctors verify students’ identities, ensure that students leave their personal belongings (phones, smart watches, “cheat sheets”) outside the testing area, provide scratch paper (since students’ papers may not enter or leave the room), and walk around the room scanning for behavior that might indicate cheating.
This unique combination of ingredients—fixed-start on-the-hour scheduling, a controlled computer environment overseen by human proctors to monitor for misbehavior and fair and time-flexible exams that thwart students sharing answers—greatly streamlines exam administration. Simply put, the effort of creating a secure exam environment and running an exam is now centralized, rather than repeated by each instructor for every exam, greatly reducing both the cost per student exam hour and the time and effort normally sunk into exam administration.
Further savings come from automated grading. Even without AI, many sophisticated questions can be automatically graded without instructor intervention because we’re capturing the students’ work in a digital form. Questions requiring numeric answers, equations, computer code and graphs can all be automatically graded, and AI offers the potential to automatically grade or at least triage an even broader range of questions, including textual short answers and hand-drawn diagrams.
When exams are cheap and easy to run, instructors can shift away from running a small number of high-stakes exams to offering frequent lower-stakes exams, which is proven to help students retain material better and reduce their test anxiety. Many instructors using our computer-based testing facilities (CBTFs) even offer students the opportunity to retake exams, incentivizing them to remediate material they hadn’t learned.
This approach is called mastery learning and was studied by noted educator Benjamin Bloom as early as 1968, but modern technology makes it practical to implement at scale without greatly increasing the amount of necessary instructor time.
For more than 10 years, CBTFs have been providing sophisticated, authentic and trustworthy testing at scale. In fall 2025, the University of Illinois’s CBTFs ran more than 130,000 exams supporting many of the largest courses on campus. The University of California, Berkeley, began piloting this approach four years ago and has now invested substantially in creating and staffing large CBTFs on its central campus, serving courses ranging in size from 30 students to more than 1,000.
Of course, we’re not suggesting that this model of assessment should replace all others: Handwritten exams, fieldwork, oral interviews and other formats may still be the best way to assess certain kinds of knowledge. But computer-based testing improves the trustworthiness and equity of many exams and streamlines operations, reducing their costs and student stress.
Blue books appeared in the U.S. in the early 20th century. Technology has come a long way since then. Computer-based testing leverages that technology (including AI) in the service of fair, unique and time-flexible exams that improve both learning outcomes and quality of life for students, while assuring academic integrity more effectively than blue books. We’re convinced that computer-based testing, thoughtfully administered and operationally efficient, is a superior path to trustworthy student-centric assessment at scale.