This is an example scorecard users can see after their conversation.
A scorecard is a structured set of Yes/No checks for a practice conversation, helping users to reflect on their behaviors demonstrated in a conversation.
Item priorities → point weights
Priority | Points |
Important | 3 |
Standard | 2 |
Low | 1 |
Not-Scored | 0 |
Use higher weights for must-do behaviors. Use Not-Scored for guidance you want visible but excluded from scoring.
How scoring works
Each item yields Yes = 1 or No = 0.
Weighted item score = Yes/No × points.
N/A items are excluded from both numerator and denominator.
Overall % = (Σ weighted Yes) ÷ (Σ points of applicable, scored items) × 100
Create a scorecard
Click Scorecards in the navigation bar → Create new. Name it and optionally link a knowledge source.
Add sections. Keep headers short (e.g., Opening, Discovery, Alignment, Closing).
Add items. One behavior per item.
Set a priority per item (Important, Standard, Low, or Not-Scored).
Save. The scorecard appears under My scorecards.
Limit: 30 items per scorecard.
Scorecard rules
Each rule exists to protect accuracy, simplicity in behavior checking, and framework/methodology adherence.
Start every item with “Did the participant…”
Why: Standardizes pattern matching, reduces ambiguity, and aligns to behavior frameworks.
If ignored: Vague prompts (“How well…”) lead to inconsistent scoring.
Good: Did the participant state the meeting purpose?
Bad: What was the purpose of the meeting?
Binary answers only (Yes/No).
Why: Removes subjective scales, simplifies automation, and improves inter-rater reliability.
If ignored: “Somewhat/partly” creates disagreement and noisy data.
Good: Did the participant confirm the decision maker?
Bad: How well did the participant identify the decision maker?
≤300 characters, one idea per item.
Why: Short checks improve LLM extraction accuracy and human scanning speed.
If ignored: Long or double-barreled items lower precision and inflate error rates.
Item must be observable in the written transcript.
Why: Text is the single source of truth. Verity does not use voice or camera for the evaluation.
If ignored: OpenAI will hallucinate about tone or body-language.
No OR/AND conditions, only one behavior per item.
Why: One behavior per check maps cleanly to frameworks and prevents false passes.
If ignored: Scores become unclear about which action earned the point.
Be descriptive yet concise.
Why: Enough specificity to find the signal; short enough to avoid drift.
If ignored: Vague items invite interpretation and reduce reliability.




