On a call, judgment is appropriately dichotomous. I had quietly imported that certainty into my classroom, where it did not belong.
The moment of learning I did not expect was the realization that my deeply ingrained field mindset, the binary pass/fail reflex, does not fit every kind of assessment in education. On a call, judgment is appropriately dichotomous: the intervention worked or it did not, and there is rarely a second attempt. That certainty felt like rigor, and I had quietly imported it into my classroom, treating each OSCE and written exam as a verdict on whether a student was or was not a paramedic, rather than as evidence about where their learning needed to go next. It is a reflex I have since begun interrupting in my own simulation debriefs.
This course unsettled that assumption. I came to see that an assessment system built entirely on isolated, high-stakes verdicts can actually undermine the learning it claims to measure. Schuwirth and van der Vleuten (2019) trace how health professions education worked its way through three phases: first treating assessment as a measurement problem, preoccupied with the reliability and validity of each individual test; then running up against the limits of competency frameworks; and finally arriving at programmatic assessment, where the value of any single instrument lies not in what it measures on its own but in the data point it contributes to a larger whole. The critical shift, as they frame it, is from trusting individual instruments to combining information across them: in a programmatic model, a weak performance on one station is not a final verdict but one source to be triangulated against the others, and every assessment, whether a formal test, a simulation, or informal feedback, becomes meaningful only when it is analyzed and translated into a learning goal. Seen through that lens, my pass/fail reflex was not rigor; it was an inherited habit that left a great deal of learning on the table.
The questions this raised were practical ones. How do I preserve the genuine, non-negotiable safety standards of paramedicine, where some errors truly are critical, while refusing to let every assessment collapse into a single verdict? How do I balance the formative space students need in order to fail safely against the summative certainty the public and the profession demand? Working through the course, I came to see that the answer was not to abandon pass/fail but to be precise about where it belongs.
Learning artifact 1
Two Standards, One Program
My error was treating two different kinds of standard as one. They are not in conflict; the fixed floor bounds the rising one.
The fixed floor: critical skills
Discrete critical skills, including airway management, CPR, spinal immobilization, and wound management, tested as isolated stations and scored pass/fail. The standard does not develop, because failure is objective and does not soften with time.
Pass / Fail, held flatCorrect in principle. Wrong only when applied to everything.
The rising floor: clinical judgment
Select a semesterTap a step on the rising floor to see how the standard converges on the certification bar as the student approaches it.
Fixed floor (does not move) Rising floor (converges on certification)
Interactive: select each semester on the rising floor.
My program, I realized, already runs on two different kinds of standard, and my error had been to treat them as one. The first is a fixed floor of discrete critical skills, airway management, cardiopulmonary resuscitation, spinal immobilization, and wound management among them, which are tested as isolated skill stations and scored pass/fail. Here the binary judgment is appropriate and does not develop, because failure is objective and the standard does not soften with time: the student chose the wrong-size oropharyngeal airway, or failed to maintain a mask seal during manual ventilation. This is the part of my field instinct that was correct all along. Pass/fail was never wrong in principle; it was wrong only when I applied it to everything.
The second standard is clinical judgment, and this is where the floor rises across the program. In first semester I extend wide latitude, tolerating delayed treatments or incorrectly identified priorities as developmental, and I prompt heavily during scenarios. In second semester the latitude narrows: students now work from medical directives, and I extend grace on the timing and recall of those directives only in situations where critical patient harm would not occur. Prompting is reduced. By third semester the standards are strict and assessed against a global rating scale, the heavy prompting of first semester has fallen to nothing, and the scenario is run in silence, requiring the student to exercise judgment alone, much as a paramedic must at three in the morning when no examiner is present. The same error that was formative in September has become disqualifying by the final semester, because the standard converges on the certification bar as the student approaches it. This is precisely the migration of the summative judgment to the trajectory that Schuwirth and van der Vleuten (2019) describe, and the deliberate withdrawal of scaffolding is the cultivation of self-directed capability that Loeng (2020) and Akyıldız (2019) argue cannot simply be assumed.
The two standards meet in second semester, and that hinge is what resolves the tension for me. The fixed floor of critical patient harm is the boundary that caps how much developmental grace I can extend on judgment: I can be generous about a mistimed directive, but not about one that would harm a patient. The rising floor and the fixed floor are therefore not in conflict; the fixed floor bounds the rising one. This alignment is also more than pedagogical. In Ontario, students are certified to perform controlled acts through a base hospital program in which standardized scenarios are scored on a global rating scale by trained raters, against a cut score, before authorization is granted. My strict, silent, GRS-scored third semester is a deliberate rehearsal of that certifying standard under the conditions the student will actually face, which means my developmental ramp is not arbitrary: it scaffolds the student toward the public-safety gate the whole program points at, then removes the scaffold just before they reach it.