Longtime readers will know that I’m a believer in the power of alternative grading (what was once called ungrading) in order to free students from some of what I see as the counterproductive structures of schooling when it comes to helping students learn to write. I know that it works because I’m convinced by the evidence shared with me by my students, but what I know is not the same as what I can prove. On Bluesky I appreciated some thoughtful threads from today’s guest, Keegan Lannon, questioning what sort of evidence we should be considering and gathering when it comes to proving the efficacy of alternative grading, so I asked him to write some of these thoughts into what I share with you all below. —JW

After years of following the discourse on Twitter/X, reading the op-eds and blogs and diving into the research around grades, it became clear to me that grades were, at best, a hindrance to education; at worst, grades are actively harming the students cognitively, emotionally and in some cases physically. I was convinced by the mountains of evidence, both qualitative and quantitative, that identified the many varied ways grades cause problems.

So, in 2023, I took the plunge and did a ground-up redesign of my trusty first-year writing syllabus, adopting aspects of a few of the various alternative grading practices people had written about under the umbrella term “ungrading.” I spent the whole summer reworking all the policies to focus on student learning and de-emphasize grades for the fall 2023 semester. I excitedly rolled it out on Blackboard and set out into a brave new world of assessment.

It did not go well.

There were a variety of issues, from the minor (needing more soft deadlines to encourage timely completion of low-stakes assignments) to the major (what knowledge of writing process students brought to class varied widely—especially among the international students). For each subsequent semester, I tinkered with the portions of my syllabus where I articulate the grading policies and practices, hopefully addressing the issues that arose over the previous semesters. With each revision, I find a nagging, intrusive thought keeps floating into my consciousness: Is this any better?

I will admit that the discourse surrounding these alternative assessment models is intoxicating in its moral clarity and confidence. It’s hard not to get swept up in the community, which is so enthusiastically looking for ways to better engage with students by building more effective learning environments.

There is a raft of studies showing the harm done to students by grades. Some of the studies show how grades can crater motivation. Some show how grades dampen confidence. Others will show how grades increase stress and can lead to mental health issues. There is no shortage of compelling data that quantifies how the variety of ways grades are harmful. It is an ethically dubious proposition that we continue to grade as we have, and there is a moral imperative for every educator to seriously consider their assessment models and make every good-faith effort to reduce harm.

That said, the ungrading movement is not without its detractors. Maggie Fernandes, Emily Brier and Megan McIntyre identify that the big tent of alternative assessment practices and the sometimes-misleading language used in the discourse can trick both students and instructors into believing they are making a change, even if they aren’t. The authors write,

“We still teach in places that assign grades. We give letter grades at the end of each semester; some of us also give multiple, mandatory mid-term letter grades to indicate a student’s progress in our course. Our students deserve to understand how grades—grades that often significantly impact their future—work in our classes. Saying we don’t care about grades, or we don’t believe in grades doesn’t seem sincere when we are still giving grades at the end of the semester. Acknowledging these facts is a vital foundation for having an honest conversation about writing assessment practices and their impact on our students.”

I, too, find myself worried that we have identified that the ladder we’ve built for our students is giving everyone splinters, but this new ladder—carefully and intentionally constructed to avoid some of the problems of the last ladder as it may be—is not yet shown to give fewer splinters. There are a couple of reasons why: First, ungrading is less a set of best practices and more a philosophical approach to assessment that de-emphasizes grades through a plethora of disparate practices. Any evaluation of a set of practices will not apply to all, or even many, of the ungrading pedagogies.

Second, classroom learning—all learning, really—is idiosyncratic and multifaceted. No two students learn in the same way or at the same pace. One student might pick up the concept fairly quickly, while another takes the whole semester; it’s hard to measure students who are heading to the same place but take very different paths to get there and arrive at different times. For some, this idiosyncratic nature of learning simply makes it too difficult to identify and control for variables that affect learning, and thus any attempt to quantify the effect of any assessment change would either be impossible or unnecessary. As John Warner writes,

“If we cannot agree on a shared measure of what success looks like, how can we come to any kind of consensus on what makes for effective teaching? Anyone who will cite scores on tests that require a five-paragraph essay as a response as a measurement of progress (or lack thereof) is in a different universe of values from me. That evidence is literally meaningless … to me.”

I don’t disagree. Learning is hard enough to even define, let alone assess. But we should not shy away from the complexity. While it might be impossible to quantify learning more broadly, we can measure and quantify some of the effects of learning.

There are some useful models that might be used to determine if some practices are more effective than traditional assessment models. Dawson R. Hancock out of the University of North Carolina at Charlotte, for example, explored how test anxiety and perceptions of the instructor could negatively impact achievement and motivation. When Hancock wrote that essay in 2010, most instructors were using the traditional letter and number grades, so it would be interesting to see if an instructor whose teaching materials clearly signaled their commitment to a set of alternative assessment practices might reduce the anxiety students feel in the classroom space and perhaps maintain (or even better, generate more) motivation and improved achievement markers.

Granted, measuring achievement on tests not only measures just a small part of the learning process, but doing so relies on traditional number and letter grading, which alternative assessment models are trying to avoid. The experiment design would need to be adapted to tie achievement less to the “correct” answers on a test and more to the standards the instructor uses to determine learning.

For example, in my first-year writing class, I want my students to approach revision not merely as a way to “correct” grammar mistakes, but as an integral part of the writing process that makes meaningful changes to the text. I could measure the complexity of the revisions (i.e., did it move past fixing commas to making global changes to the text). Better still, I could measure their willingness or attitudes toward making those changes using some version of Thurstone and/or Likert attitude scales. With more narrative data, such as free-response survey questions, a sentiment analysis algorithm could be used to determine what attitudes undergird their responses.

I’m not naïve. Certainly, identifying which assessment approaches avoid the same anxiety-inducing tendencies of more traditional assessment models gives us limited information. There are still dozens of confounding variables to consider. It could be possible that it was the instructor that influenced the student’s attitudes and reduced anxiety more than any one policy change. It could be the students in class were more amenable to alternative models than other student demographics. Still, any data is useful, and whatever can be collected, qualified as it may be, is still a step toward knowing that we aren’t harming our students. Those steps seem like worthwhile steps to take. 

Keegan Lannon is a lecturer in the English Department at the University of Illinois at Chicago, teaching classes in first-year writing, rhetoric and comics. You can find him on Bluesky, where he posts about writing, teaching, comics and his lifelong love of the Cubs.

Next Story

Written By

Share This Article

More from Just Visiting