Skip to main content

How to Write Great Scorecards

Written by Richard Lombos

1. Start Every Item with “Did the participant…”

Meaning:

Every item should begin with “Did the participant…”, so each is clear and direct about the observed behavior.

Why it’s important:

This phrasing keeps questions focused on observable actions and makes reviewing/automating scoring easier. LLMs look for written evidence matching the prompt.

What goes wrong if not followed:

Questions may become ambiguous (“Was the call productive?”), subjective (“How well did the rep do?”), or confusing for both the AI and humans. LLMs might not know what to look for.

Scenario

Good Example (Yes/No, text-based)

Bad Example (ambiguous)

Sales

Did the participant confirm the customer’s goal?

What was the customer’s main goal?

Sales

Did the participant share next steps at the end?

How clear were the next steps?

Leadership

Did the participant ask for feedback clarification?

Did the participant seem open to feedback?

Leadership

Did the participant state a specific improvement action?

Was the participant committed to change?


2. Binary Answers Only (Yes/No)

Meaning:

Every item must be answerable with “Yes” or “No” based only on what’s written in the transcript.

Why it’s important:

LLMs cannot make reliable judgments on scales (“1-5,” “poor-good”). Binary questions avoid subjective scoring and are easier to automate.

What goes wrong if not followed:

Reviewers (AI or human) might disagree about what “partly” or “somewhat” means. This causes inconsistent, unreliable feedback.

Scenario

Good Example (Yes/No)

Bad Example (scale or subjective)

Sales

Did the participant offer a product demo?

How well did the participant demonstrate the product?

Sales

Did the participant ask about decision criteria?

Did the participant effectively uncover decision criteria?

Leadership

Did the participant acknowledge impact?

Did the participant communicate impact well?

Leadership

Did the participant propose a clear next step?

Was the participant proactive?


3. ≤300 characters, One Idea Per Item

Meaning:

Each scorecard item must be concise (≤300 characters) and target only one observable action or behavior.

Why it’s important:

LLMs handle short, focused prompts best. Long, complex questions (or combined checks) reduce accuracy and can be misunderstood.

What goes wrong if not followed:

Overly long or multi-part questions confuse the model and reviewers. Important points might be missed or marked incorrectly.

Scenario

Good Example (short, single idea)

Bad Example (long, combined ideas)

Sales

Did the participant ask about customer pain points?

Did the participant uncover pain points and offer a solution in the same question?

Sales

Did the participant confirm the agenda early in the call?

Did the participant confirm the agenda or move the conversation forward efficiently?

Leadership

Did the participant request specific feedback?

Did the participant request feedback or share their own ideas for improvement?

Leadership

Did the participant summarize next steps?

Did the participant summarize next steps and express appreciation for the conversation?


4. Make It Observable—Written Transcript Only

Meaning:

Write items that are provable in the written transcript. Avoid tone, body language, or anything not directly spoken/written.

Why it’s important:

LLMs have no access to audio, tone, or facial expressions—only the exact words. Scoring must be 100% based on visible, written actions.

What goes wrong if not followed:

Items like “Did the participant sound confident?” cannot be scored by LLMs and will lead to errors, inconsistent scores, or false assumptions.

Scenario

Good Example (observable)

Bad Example (not observable)

Sales

Did the participant ask at least one open-ended question?

Did the participant sound interested?

Sales

Did the participant confirm the other person’s understanding?

Did the participant use an enthusiastic tone?

Leadership

Did the participant provide a date or example in feedback?

Did the participant seem empathetic?

Leadership

Did the participant ask for clarification when feedback was vague?

Did the participant have positive body language?


5. No “OR” Conditions (One Condition Per Item)

Meaning:

Scorecard items should only check for a single, specific action—not “X or Y” or “X and Y.” Each “OR” should become its own Yes/No item.

Why it’s important:

If you use “OR,” it’s unclear which behavior earned the score. LLMs can only reliably find one action per prompt.

What goes wrong if not followed:

You get unreliable results: a call may get a Yes for only doing the easier half, or fail because reviewers disagree on which action counts.

Scenario

Good Example (one condition)

Bad Example (OR or multi-condition)

Sales

Did the participant share the call agenda early?

Did the participant share the agenda or summarize next steps?

Sales

Did the participant provide at least one proof point?

Did the participant share a case study or talk about results or value?

Leadership

Did the participant ask for clarification?

Did the participant ask for clarification or propose next steps?

Leadership

Did the participant acknowledge impact?

Did the participant mention team or customer impact or discuss improvement?


6. Mark Optional/Conditional Items as Non-Mandatory (Allow N/A)

Meaning:

For items that don’t apply to every call (like objection handling), make them non-mandatory so “N/A” can be scored.

Why it’s important:

AI shouldn’t penalize a rep for not handling objections if there were none. N/A makes the score fair and relevant.

What goes wrong if not followed:

Mandatory conditional items penalize for things that never happened, leading to unfair scores and demotivated reps.

Scenario

Good Example (Non-Mandatory)

Bad Example (Mandatory, penalizing)

Sales

If an objection occurred, did the participant address it?

Did the participant handle objections? (counts as No even if no objection)

Sales

If pricing was discussed, did the participant confirm budget?

Did the participant discuss budget? (always required)

Leadership

If disagreement happened, did the participant respond respectfully?

Did the participant manage conflict well? (even if no conflict)

Leadership

If feedback was unclear, did the participant ask for clarification?

Did the participant clarify feedback? (even if not needed)


7. Be Descriptive yet Concise

Meaning:

Each item should be detailed enough to avoid guessing, but short enough to read quickly.

Why it’s important:

LLMs (and human reviewers) work best with clear, targeted prompts that say exactly what to look for—nothing more, nothing less.

What goes wrong if not followed:

If the item is too vague or too long, AI (and humans) may guess, disagree, or miss important context.

Scenario

Good Example (descriptive + concise)

Bad Example (vague or wordy)

Sales

Did the participant ask open-ended questions to learn about needs?

Was the participant inquisitive?

Sales

Did the participant summarize agreed next steps at the end?

Did the participant end the call in a positive way and with clear follow-up?

Leadership

Did the participant describe a specific example in their feedback?

Did the participant give helpful feedback?

Leadership

Did the participant propose at least one specific improvement?

Did the participant seem eager to improve and ready to act?


8. Ensure Consistency

Meaning:

Use the same structure, wording style, and answer format throughout your scorecard.

Why it’s important:

LLMs are more accurate and humans make fewer mistakes when all items “look and feel” the same. This helps with batch scoring and trend analysis.

What goes wrong if not followed:

Inconsistent format creates confusion, errors, and inconsistent reporting.

Scenario

Good Example (consistent structure)

Bad Example (mixed styles)

Sales

Did the participant ask about decision criteria?

How did the rep approach decision criteria?

Sales

Did the participant avoid filler words?

Was language clear, or did the rep use lots of filler?

Leadership

Did the participant invite the other person’s perspective?

Was the other person invited to share, or did you feel they were included?

Leadership

Did the participant state a clear action item?

Did the feedback feel actionable or concrete?


9. Prioritize Objectivity

Meaning:

Write questions about things that can be proven in the transcript, not about feelings or opinions.

Why it’s important:

LLMs can’t “feel” if something was good or not; they can only see what was said or written. Objectivity keeps scoring fair and repeatable.

What goes wrong if not followed:

Subjective items (“Did the rep inspire confidence?”) create bias, disagreement, and poor feedback for learners.

Scenario

Good Example (objective)

Bad Example (subjective)

Sales

Did the participant list at least three product features?

Did the participant present the product effectively?

Sales

Did the participant reference the customer’s industry?

Did the participant show good industry knowledge?

Leadership

Did the participant request a concrete next step?

Did the participant motivate the team well?

Leadership

Did the participant share an example with a date or timeframe?

Did the participant create urgency and excitement?


10. Iterate and Refine

Meaning:

Treat your scorecard as a living document. Test it, gather feedback, and keep improving each item.

Why it’s important:

Business needs, AI accuracy, and best practices change over time. Regular updates keep your scorecard relevant and fair.

What goes wrong if not followed:

Scorecards become outdated, unfair, or too easy/hard—leading to missed opportunities for real improvement.

Scenario

Good Example (added/refined after review)

Bad Example (never updated, unclear)

Sales

Did the participant confirm understanding after summarizing?

Did the participant keep alignment throughout the call or adapt as needed?

Sales

Did the participant provide a proof point when recommending a solution?

Did the participant persuade the other person?

Leadership

Did the participant adjust based on feedback trends from previous sessions?

Did the participant maintain consistency across different meetings?

Leadership

Did the participant close the meeting with an agreed next step?

Did the participant create a good ending to the discussion?

Did this answer your question?