The methodology

How we make assessments you can stand behind.

Most tools go straight from your material to questions and leave you to catch the problems. SayaLab works on both sides of generation: it reads your material first to see what it can fairly test, then validates every question it writes and shows you the evidence. Review becomes confirming, not catching.

Why generation isn't enough

Generating questions is the easy part.

A model can write a hundred questions in seconds. The hard part is whether they test the right level of thinking, map to your objectives, and have wrong answers that are wrong for the right reasons. The research is consistent: left unchecked, AI-generated questions skew toward recall and carry item-writing flaws. Roughly 30 to 43% get rewritten, and 23 to 38% carry item-writing flaws in the first place.

AI questions skew to recall, measurably weaker at higher Bloom's levels.PLOS ONE 2024 · BMC 2025
Distractors are often implausible, the most common item-writing flaw.PMC 2025
Done right, validated AI assessments can match human-authored quality. The ceiling is higher than the default, which is why the method matters: analyze before, validate after.Nature npj 2026 · Stanford 2025
Before generation

It reads your material first.

Instructional design has always required an analysis step before you write a single question. SayaLab automates it: before generating anything, it works out what your material can fairly teach, and tells you where it can't.

  • Learning objectivespulled from your material in measurable terms, and yours to edit, add to, or remove before anything is generated.
  • The Bloom's ceilingthe highest level of thinking each part of your material can actually support, with the reason why.
  • Content sufficiencyeach objective marked sufficient, thin, or insufficient for the level you want. If the material is too thin, it says so rather than inventing a question.
  • Answerability and audience fitwhether each objective can be tested from the material alone, and whether its depth matches the audience you are teaching.
After generation

It validates every question, independently.

Each question runs a validation pass before you ever see it, a separate check rather than the model grading its own homework.

  • Bloom's level, re-classified blinda separate pass classifies the question's actual level without seeing the target. The level is a result we verify, not a label you picked.
  • Objective alignmentthe question genuinely measures its objective: a learner who has met the objective can answer it, and one who has not reliably cannot.
  • Distractor plausibilityfor multiple-choice, every wrong answer is a real misconception, not a giveaway or filler.
  • Source fidelitythe question is answerable from your material; anything that would need outside knowledge is flagged.
  • Clarityunambiguous wording, a single correct answer, no double negatives, vocabulary that fits the audience.

Questions that fall short are regenerated automatically, up to two tries. If one still cannot clear the bar, it reaches you flagged, with the reason, so you decide.

The evidence

Every question shows its work.

You don't get one score for the whole assessment. Each question carries its own evidence: the Bloom's level we verified, the objective it maps to, how plausible its distractors are, and whether it is grounded in your source, with a quality score from 0 to 1. That is what makes review fast: you are confirming the evidence, not hunting for problems.

Our bar

The standard every assessment is held to.

These are the targets our validation is built to meet, the bar a question has to clear to reach you unflagged. They are our internal standard, not a measured marketing claim; as our validation set grows, we will publish results.

≥ 85%

questions matching their intended Bloom's level on blind re-classification

≥ 80%

questions that validly measure their stated learning objective

≥ 90%

multiple-choice distractors rated plausible

≥ 95%

questions answerable from your own source material

What it's built on

Established pedagogy, not invented rules.

  • Bloom's taxonomythe six cognitive levels we classify against: remember, understand, apply, analyze, evaluate, create.
  • Biggs' Constructive Alignmentobjectives, questions, and evidence are deliberately lined up, so what you assess is what you set out to teach.
  • Evidence-Centered Designevery question is built to elicit evidence of a specific competency, not just surface recall.
  • Item-writing researchdistractor quality is scored against the published literature on what makes a wrong answer a real misconception.
Questions

Common questions.

Is SayaLab a quiz generator?
It generates assessments, but generation is the smallest part. Most tools go straight from your text to questions and leave you to catch the problems. SayaLab reads your material first to see what it can fairly test, validates every question after it is written, and shows you the evidence, so review means confirming, not catching.
How is each question's Bloom's level decided?
Not by a label you pick. After a question is written, a separate pass re-classifies its Bloom's level blind, without seeing the target level. The level is a result we verify, not a setting you choose.
What happens if my material can't support a question?
Before generating, SayaLab marks each objective as sufficient, thin, or insufficient for the cognitive level you want. If the material is too thin to test something well, it tells you instead of inventing a question anyway.
Can it invent answers or rely on outside knowledge?
Every question is checked for source fidelity: it must be answerable from the material you uploaded. Questions that would require outside knowledge are flagged rather than shipped silently.
Can AI-generated assessments really match human-written ones?
With the right method, the research says yes. Peer-reviewed studies find that validated, refined AI questions can match human-authored quality. That is the case for validating every question rather than handing you raw output to fix.

Generate the assessment. Keep the pedagogy.

Get started free

Free tier, no card required.