toolkit / stream-exam-engine

Built on PLEW — patented · 10,000+ questions produced

Questions to examination standard. Then solved.

Exam Question Engine generates reading questions and passages to examination standard, and solves everything it writes. It produces to the formats of the 수능 and 내신, and of TARA and the LNAT; every question it writes ships with a full solution — vocabulary, grammar, logical flow, simplification, and the strategy by which the question is efficiently answered.

Enter the engine Browse the question sets

10,000+

QUESTIONS PRODUCED

22

DIFFICULTY INDICATORS

4.8%

RMSE VS PUBLISHED RATES

STREAM EXAM ENGINE — DEMONSTRATION

function: F1 — question matching

Illustrative demonstration. The live demo connects to the engine in production.

The method

PLEW, the method underneath

The engine is built on PLEW, a reading method developed in our classrooms in 2019 and since patented. PLEW extends the Toulmin model of argument to the examined passage: every passage is decomposed into four structural elements. The Purpose sentence carries the main claim. The Logic is the inferential structure connecting claims. The Evidence is classified by type and by its logical relationship to the claim. The Weaknesses are the concessions and qualifications the writer admits. A student trained on PLEW reads an unfamiliar passage the way an examiner reads it — knowing what each sentence is for. The same analysis, run in software, is what lets the engine write questions that behave like real ones.

P — PurposeThe main claim

The assumption that transparency alone disciplines institutions has a mixed record.

E — EvidenceSupport, and its relation to the claim

Mandatory disclosure in several regulated industries produced fuller reporting without a measurable change in conduct.

L — LogicThe inferential structure

Where audiences lack the means to act on what they learn, disclosure substitutes for reform rather than producing it.

W — WeaknessConcessions the writer admits

Transparency does, admittedly, alter behaviour where an attentive audience holds real power to sanction.

Illustrative passage, annotated as the engine reads it. PLEW — patented, developed 2019.

The decomposition is not an exercise in labelling. Every examined question type tests one of the four elements, and the mapping is exact:

요지 · 주장Main idea and claim

Tests P — identification of the purpose sentence.

빈칸 · 순서Gap-fill and sentence ordering

Tests L — reconstruction of the logic.

어휘Vocabulary in context

Tests E — comprehension of the evidence structure.

함의추론Implied meaning

Tests L — inference of logic the passage leaves unstated.

PLEW taught in the classroom — students working from the whiteboard
The method before the software. PLEW began as classroom teaching at the Gangnam campus.

PLEW is in commercial use. Korean education companies are building integrations on it; the most recent delivery is the solutions sheets produced for A Dot's Dshare platform.

The indicators

What makes a question hard, measured

PLEW is the first of two layers. The second is a set of reasoning indicators derived from the Thinking Skills Assessment — the reasoning test developed by Cambridge Assessment for Oxford admissions — and adapted for examined English. Six operations are scored on every passage and every question: assumption identification, stated and unstated; parallel-structure recognition; the distinction between cause and correlation; counter-argument handling; analogy evaluation; and the logical distance between evidence and claim.

In full, the difficulty model runs on twenty-two indicators across two dimensions, passage and question, combined into a single predicted answer rate. On the passage side: sentence complexity, vocabulary tier, logical chain length, idea density, assumption depth, and whether the main idea is stated, implied, or distributed across the passage. On the question side: distractor design, the number of inference steps between passage and correct answer, and the paraphrase distance between the passage's language and the language of the options.

Discourse type
Persuasive — claim, evidence, concession
Main idea
Stated, opening position
Assumption depth
2 of 3 — intermediate; the sanctioning-power premise is unstated
Logical chain
3 inferential steps from evidence to claim
Concession
Present — final sentence
Predicted 정답률
46% if set as 함의추론 at this reading load

The illustrative passage above, profiled on six of the twenty-two indicators.

The taxonomy is published; the calibration is proprietary. The weight each indicator carries was not set theoretically — the weights are derived from the responses of the TARA research cohort, more than 5,000 students assessed on the same reasoning framework. Cambridge Assessment and Pearson publish their construct definitions and withhold their item parameters. The engine operates on the same principle.

The functions

The functions, as a specification

The engine's format descriptions follow the convention of the tests themselves.

F1
Takes a given passage and writes questions of any set type against it, to examination standard.
F2
Produces new passages at a chosen difficulty, from real sources.
L1–L3
Rewrites any passage across three logical-distance levels, holding the claim constant.
Solutions
Every output carries its solution set: vocabulary, grammar, logical flow, simplification, strategy.

Coverage: 수능 and 내신 English, TARA format, LNAT format. 10,000+ questions produced to date.

L1 · Lexical-syntactic

Every element of the inference chain is preserved. The variation is in vocabulary and syntax only.

L2 · Structural

The core inference chain is preserved; sub-arguments may be reordered or consolidated.

L3 · Architectural

Only the main claim and its logical type survive. The reasoning architecture is rebuilt around them.

The paraphrase levels are defined by logical distance — whether the inference chain survives the rewrite — not by how different the surface text looks.

The measure

4.8%

The engine's difficulty predictions are tested against the strongest public standard available: the published answer rates of the 수능 — the percentage of candidates nationally who answer each question correctly, released after every examination. On a validation set of 140 questions drawn from the public historical papers, the engine's predicted answer rates match the published national rates to within 4.8 per cent root-mean-square error. The prediction is made before the published rate is consulted; the test is falsifiable, and anyone with the public papers can re-run it.

Using it

For students. For companies.

Practice

Students practise on the engine through question sets arranged by examination and type — for the 수능, by the full 유형 range including 빈칸, 순서 and 삽입 — each set with complete solutions. For LNAT and TARA users, questions can be set by type and subject area.

수능 · 내신 · TARA · LNAT

Licensing

Korean education companies are building integrations on PLEW; the most recent delivery is the solutions sheets for A Dot's Dshare. For company licensing enquiries: team@examrizz.com

For full builds, see Stream Consulting →