toolkit / stream-exam-engine
Built on PLEW — patented · 10,000+ questions produced
Exam Question Engine generates reading questions and passages to examination standard, and solves everything it writes. It produces to the formats of the 수능 and 내신, and of TARA and the LNAT; every question it writes ships with a full solution — vocabulary, grammar, logical flow, simplification, and the strategy by which the question is efficiently answered.
Illustrative demonstration. The live demo connects to the engine in production.
The method
The engine is built on PLEW, a reading method developed in our classrooms in 2019 and since patented. PLEW extends the Toulmin model of argument to the examined passage: every passage is decomposed into four structural elements. The Purpose sentence carries the main claim. The Logic is the inferential structure connecting claims. The Evidence is classified by type and by its logical relationship to the claim. The Weaknesses are the concessions and qualifications the writer admits. A student trained on PLEW reads an unfamiliar passage the way an examiner reads it — knowing what each sentence is for. The same analysis, run in software, is what lets the engine write questions that behave like real ones.
The decomposition is not an exercise in labelling. Every examined question type tests one of the four elements, and the mapping is exact:
PLEW is in commercial use. Korean education companies are building integrations on it; the most recent delivery is the solutions sheets produced for A Dot's Dshare platform.
The indicators
PLEW is the first of two layers. The second is a set of reasoning indicators derived from the Thinking Skills Assessment — the reasoning test developed by Cambridge Assessment for Oxford admissions — and adapted for examined English. Six operations are scored on every passage and every question: assumption identification, stated and unstated; parallel-structure recognition; the distinction between cause and correlation; counter-argument handling; analogy evaluation; and the logical distance between evidence and claim.
In full, the difficulty model runs on twenty-two indicators across two dimensions, passage and question, combined into a single predicted answer rate. On the passage side: sentence complexity, vocabulary tier, logical chain length, idea density, assumption depth, and whether the main idea is stated, implied, or distributed across the passage. On the question side: distractor design, the number of inference steps between passage and correct answer, and the paraphrase distance between the passage's language and the language of the options.
The illustrative passage above, profiled on six of the twenty-two indicators.
The taxonomy is published; the calibration is proprietary. The weight each indicator carries was not set theoretically — the weights are derived from the responses of the TARA research cohort, more than 5,000 students assessed on the same reasoning framework. Cambridge Assessment and Pearson publish their construct definitions and withhold their item parameters. The engine operates on the same principle.
The functions
The engine's format descriptions follow the convention of the tests themselves.
Coverage: 수능 and 내신 English, TARA format, LNAT format. 10,000+ questions produced to date.
Every element of the inference chain is preserved. The variation is in vocabulary and syntax only.
The core inference chain is preserved; sub-arguments may be reordered or consolidated.
Only the main claim and its logical type survive. The reasoning architecture is rebuilt around them.
The paraphrase levels are defined by logical distance — whether the inference chain survives the rewrite — not by how different the surface text looks.
The measure
4.8%
The engine's difficulty predictions are tested against the strongest public standard available: the published answer rates of the 수능 — the percentage of candidates nationally who answer each question correctly, released after every examination. On a validation set of 140 questions drawn from the public historical papers, the engine's predicted answer rates match the published national rates to within 4.8 per cent root-mean-square error. The prediction is made before the published rate is consulted; the test is falsifiable, and anyone with the public papers can re-run it.
Using it
Students practise on the engine through question sets arranged by examination and type — for the 수능, by the full 유형 range including 빈칸, 순서 and 삽입 — each set with complete solutions. For LNAT and TARA users, questions can be set by type and subject area.
수능 · 내신 · TARA · LNAT
Korean education companies are building integrations on PLEW; the most recent delivery is the solutions sheets for A Dot's Dshare. For company licensing enquiries: team@examrizz.com
For full builds, see Stream Consulting →