How we make assessments you can stand behind.
Most tools go straight from your material to questions and leave you to catch the problems. SayaLab works on both sides of generation: it reads your material first to see what it can fairly test, then validates every question it writes and shows you the evidence. Review becomes confirming, not catching.
Generating questions is the easy part.
A model can write a hundred questions in seconds. The hard part is whether they test the right level of thinking, map to your objectives, and have wrong answers that are wrong for the right reasons. The research is consistent: left unchecked, AI-generated questions skew toward recall and carry item-writing flaws. Roughly 30 to 43% get rewritten, and 23 to 38% carry item-writing flaws in the first place.
It reads your material first.
Instructional design has always required an analysis step before you write a single question. SayaLab automates it: before generating anything, it works out what your material can fairly teach, and tells you where it can't.
- Learning objectives — pulled from your material in measurable terms, and yours to edit, add to, or remove before anything is generated.
- The Bloom's ceiling — the highest level of thinking each part of your material can actually support, with the reason why.
- Content sufficiency — each objective marked sufficient, thin, or insufficient for the level you want. If the material is too thin, it says so rather than inventing a question.
- Answerability and audience fit — whether each objective can be tested from the material alone, and whether its depth matches the audience you are teaching.
It validates every question, independently.
Each question runs a validation pass before you ever see it, a separate check rather than the model grading its own homework.
- Bloom's level, re-classified blind — a separate pass classifies the question's actual level without seeing the target. The level is a result we verify, not a label you picked.
- Objective alignment — the question genuinely measures its objective: a learner who has met the objective can answer it, and one who has not reliably cannot.
- Distractor plausibility — for multiple-choice, every wrong answer is a real misconception, not a giveaway or filler.
- Source fidelity — the question is answerable from your material; anything that would need outside knowledge is flagged.
- Clarity — unambiguous wording, a single correct answer, no double negatives, vocabulary that fits the audience.
Questions that fall short are regenerated automatically, up to two tries. If one still cannot clear the bar, it reaches you flagged, with the reason, so you decide.
Every question shows its work.
You don't get one score for the whole assessment. Each question carries its own evidence: the Bloom's level we verified, the objective it maps to, how plausible its distractors are, and whether it is grounded in your source, with a quality score from 0 to 1. That is what makes review fast: you are confirming the evidence, not hunting for problems.
The standard every assessment is held to.
These are the targets our validation is built to meet, the bar a question has to clear to reach you unflagged. They are our internal standard, not a measured marketing claim; as our validation set grows, we will publish results.
questions matching their intended Bloom's level on blind re-classification
questions that validly measure their stated learning objective
multiple-choice distractors rated plausible
questions answerable from your own source material
Established pedagogy, not invented rules.
- Bloom's taxonomy — the six cognitive levels we classify against: remember, understand, apply, analyze, evaluate, create.
- Biggs' Constructive Alignment — objectives, questions, and evidence are deliberately lined up, so what you assess is what you set out to teach.
- Evidence-Centered Design — every question is built to elicit evidence of a specific competency, not just surface recall.
- Item-writing research — distractor quality is scored against the published literature on what makes a wrong answer a real misconception.
Common questions.
- Is SayaLab a quiz generator?
- It generates assessments, but generation is the smallest part. Most tools go straight from your text to questions and leave you to catch the problems. SayaLab reads your material first to see what it can fairly test, validates every question after it is written, and shows you the evidence, so review means confirming, not catching.
- How is each question's Bloom's level decided?
- Not by a label you pick. After a question is written, a separate pass re-classifies its Bloom's level blind, without seeing the target level. The level is a result we verify, not a setting you choose.
- What happens if my material can't support a question?
- Before generating, SayaLab marks each objective as sufficient, thin, or insufficient for the cognitive level you want. If the material is too thin to test something well, it tells you instead of inventing a question anyway.
- Can it invent answers or rely on outside knowledge?
- Every question is checked for source fidelity: it must be answerable from the material you uploaded. Questions that would require outside knowledge are flagged rather than shipped silently.
- Can AI-generated assessments really match human-written ones?
- With the right method, the research says yes. Peer-reviewed studies find that validated, refined AI questions can match human-authored quality. That is the case for validating every question rather than handing you raw output to fix.