Skip to main content

How Scorecards work in Verity

Overview, Creation, Weights, and Rules

Written by Richard Lombos

This is an example scorecard users can see after their conversation.

A scorecard is a structured set of Yes/No checks for a practice conversation, helping users to reflect on their behaviors demonstrated in a conversation.


Item priorities → point weights

Priority

Points

Important

3

Standard

2

Low

1

Not-Scored

0

Use higher weights for must-do behaviors. Use Not-Scored for guidance you want visible but excluded from scoring.


How scoring works

  • Each item yields Yes = 1 or No = 0.

  • Weighted item score = Yes/No × points.

  • N/A items are excluded from both numerator and denominator.

Overall % = (Σ weighted Yes) ÷ (Σ points of applicable, scored items) × 100


Create a scorecard

  1. Click Scorecards in the navigation bar → Create new. Name it and optionally link a knowledge source.

  2. Add sections. Keep headers short (e.g., Opening, Discovery, Alignment, Closing).

  3. Add items. One behavior per item.

  4. Set a priority per item (Important, Standard, Low, or Not-Scored).

  5. Save. The scorecard appears under My scorecards.

Limit: 30 items per scorecard.


Scorecard rules

Each rule exists to protect accuracy, simplicity in behavior checking, and framework/methodology adherence.

Start every item with “Did the participant…”

  • Why: Standardizes pattern matching, reduces ambiguity, and aligns to behavior frameworks.

  • If ignored: Vague prompts (“How well…”) lead to inconsistent scoring.

  • Good: Did the participant state the meeting purpose?

  • Bad: What was the purpose of the meeting?

Binary answers only (Yes/No).

  • Why: Removes subjective scales, simplifies automation, and improves inter-rater reliability.

  • If ignored: “Somewhat/partly” creates disagreement and noisy data.

  • Good: Did the participant confirm the decision maker?

  • Bad: How well did the participant identify the decision maker?

≤300 characters, one idea per item.

  • Why: Short checks improve LLM extraction accuracy and human scanning speed.

  • If ignored: Long or double-barreled items lower precision and inflate error rates.

Item must be observable in the written transcript.

  • Why: Text is the single source of truth. Verity does not use voice or camera for the evaluation.

  • If ignored: OpenAI will hallucinate about tone or body-language.

No OR/AND conditions, only one behavior per item.

  • Why: One behavior per check maps cleanly to frameworks and prevents false passes.

  • If ignored: Scores become unclear about which action earned the point.

Be descriptive yet concise.

  • Why: Enough specificity to find the signal; short enough to avoid drift.

  • If ignored: Vague items invite interpretation and reduce reliability.

Did this answer your question?