top of page

How should you grade a marketing simulation?

Writer: Clark Boyd
Clark Boyd
2 days ago
7 min read

Updated: 24 hours ago

It is the question that comes up at every conference booth, in the last five minutes of every demo, and on almost every onboarding call once term has started. How should you grade a marketing simulation? Professors ask it more often than they ask about price, and I understand why. A simulation produces a number, the number looks like a grade, and the pull to move it straight into the gradebook is strong. It is question six in our seven questions to ask before you choose a marketing simulation.


The short answer: the score is a component, not the grade. In the courses I see working, it carries somewhere between 20% and 40%, and the rest of the marks sit on a written or spoken defence of the decisions that produced it. The score says how the student did. The defence says whether they understood why. The rest of this piece is the case for that split, and what the defence looks like in practice.


What does the simulation score actually measure?


  1. Judgement: The student read the market, allocated the budget, chose the audience, and the result reflects, in part, whether those were good choices.

  2. Luck: A simulation that models a market honestly is noisy, because real markets are. Strip the noise out and you have a spreadsheet that rewards clicking, not a market. So the noise stays, and lands in the score.

  3. The student’s ability to reverse-engineer the model. Give the score enough weight and a certain kind of student stops asking what would work in this market and starts asking what the software rewards. They will find it. The score goes up. The learning does not.


Only the first is something you set out to teach. But the score bundles all three into one figure, and once the round has closed there is no way to tell how much of it was judgement and how much was the other two.


Where does grading the score fail?


There is a name for that mistake, and a literature behind it. Baron and Hershey (1988), in the Journal of Personality and Social Psychology, called it outcome bias: people rate a decision as better when it happened to work out, even when told the outcome was partly chance, and even after agreeing that outcomes should not count. A pre-registered replication by Aiyer and colleagues (2023), in the International Review of Social Psychology, found the effect again and larger. If people who have just said outcomes should not count cannot hold the line, a professor marking eighty submissions on a Sunday will not. Put the whole grade on the score and you have built the bias into the rubric.


Bar chart comparing effect sizes for outcome bias: Baron and Hershey 1988, d 0.21 to 0.53, and Aiyer et al. 2023 replication, d 0.77 to 1.10.

The simulation literature has said something adjacent for thirty years. Anderson and Lawton (1992), in Simulation & Gaming, tested the link directly and found no significant relationship between simulation financial performance and most of their measures of learning. Gosen and Washbush (2004), reviewing decades of scholarship for the same journal, put it plainly: performance in a simulation should not be used as a proxy for learning. The score and the learning are different things.


So how should the grade be built?


Two parts.


One. The score, as a minority share. Somewhere in the 20 to 40% range is where most professors land. It keeps students honest about the outcome without making the outcome the point. Use one number. If the platform gives several metrics, pick one or combine them, because a results screen with six numbers on it invites "but I got more revenue than they did", and you will hear it in office hours until Christmas. The simulation supplies the score and you set the boundary for an A. Where the LMS integration is in place, the score posts straight into your gradebook, so this part should cost you no time at all.


Two. A written defence of the decisions. This is where most of the marks sit. The student argues what they knew at the time and why it justified the call. A defence of a decision that turned out badly, argued well, can and should outscore a win the student cannot explain. Tell them this on day one. Once a class knows reasoning is marked as heavily as outcome, they start writing down why they chose things as they go, which is the habit you were trying to build anyway.


The defence is also the part your accreditor can use. AACSB's assurance of learning standard asks for direct measures of student learning, and a marked defence mapped to an outcome is one. A score is not. And it is the debrief made assessable. Crookall (2010), in Simulation & Gaming, argued that the game generates the experience and the debriefing turns it into learning. Fanning and Gaba (2007), in Simulation in Healthcare, found the debrief was the factor most associated with what participants took away. The debrief is where many simulation programmes quietly fail, a point I made in the case method comparison, and writing it down is the cheapest fix I know. Vos (2015), in the International Journal of Management Education, surveyed educators on how they assess simulation learning and found most already lean this way: formative feedback, transparent criteria, required reflection.


What does the defence look like in practice?


Four forms, which can be combined. They share one design principle: the student’s expectation goes on the record before the result exists.


A decision log, kept across the term. Every material decision logged against the same fields: what they were trying to achieve, what evidence they had, what alternatives they considered, what they chose, what they expected to happen, what actually happened, and why. Graded on strategic coherence, use of evidence, quality of the trade-offs, and how well the gap between expectation and outcome gets explained. It also holds up against AI-written coursework better than an essay does, because the expected result has to be written down before the actual one exists, and no tool can backfill that convincingly.


A short reflection memo after each round. A few hundred words, due as the round closes, defending one decision made in it. This works because the round produces facts the student did not choose, so there is something real for their reasoning to be tested against.


Flow diagram: the plan, the round, the reflection, then the next round. The plan and reflection carry most of the marks, and the round produces the score, which takes a minority share.

Peer assessment of the defence. Students present their strategy to classmates who played the same scenario and assess each other’s reasoning directly. A classmate who faced the same market can spot a rationalised decision faster than a marker reading it cold, and it forces students to explain their choices out loud, not only on paper.


A capstone with a short oral defence. A written strategy, then a few minutes of live questioning without AI assistance, drawn from a prepared question bank. Which assumption in your plan carries the most risk. What would make you reverse this recommendation. The written work can be produced with help. The defence cannot.


What a strong submission looks like, in any of these forms:

  • A diagnosis built from the evidence the simulation gave them, not a principle recalled from a lecture.

  • A decision that follows from the diagnosis rather than one they would have made anyway.

  • An expectation stated in advance and then compared with the result.

  • A change for the next round that follows from the reason given for the gap.


Marketing simulation grading rubric: a four-by-four table with criteria Diagnosis, Decision, Expectation and Change, scored 1 to 4.

What about the leaderboard?


Keep it. Students often enjoy it, and I am not going to strip out the part of the exercise a cohort finds particularly fun. But the leaderboard is not the score. The score component is a number against a boundary you set. A leaderboard position is a rank against classmates, and the moment rank carries marks, it stops being sport and becomes an incentive pointing at the model rather than the market. Rank for fun. Score for the minority share. Defence for the rest.


What are the honest caveats?


First, the grading load is real. Reading a log or a memo for eighty students, round after round, is slower than sorting a column of scores, and a rubric makes it faster and fairer without making it fast. Anyone who tells you assessment by defence is free has not run it on a large cohort.


Second, the 20 to 40% range is what I see working. It is not a studied optimum. The literature is clear that the score is a poor proxy for learning and thin on how much of the grade it should carry, so treat the range as a starting point and move it once you have seen a cohort through.


Third, the oral defence and peer assessment do not scale evenly. A viva for thirty MBAs is an afternoon. For 180 undergraduates it is a week, and peer marking on that scale can reward the confident student over the careful one unless the criteria are tight. The log and the memo carry the weight in a large class. The viva is for where you can afford it.


I would rather have this conversation than the one about which platform has the most AI features. The grading question is the one that decides whether the simulation taught anything.


What does this mean for the professor in week four?


  1. If the score is already promised as the whole simulation grade, add one reflection memo after the next round and mark it. Tell the class why. You will learn more from twenty of those than from the whole leaderboard.


  1. If you are setting the module up now, write the split into the syllabus: 20 to 40 per cent on the score, the rest on the defence, and say on day one that a well-argued failure beats an unexplained win.


  1. If you have a capstone, add ten minutes of live questioning without AI. It is the only piece of assessment left that cannot be produced with help.


This is the shape we build our own assessment around. To talk it through against your module, book a demo. And if you grade yours differently and it works, I would like to hear it. I am not convinced there is one right split, only that a hundred per cent on the score is the wrong one.


Comments


bottom of page