STUD
← Plays

Build · Create · Analysis & Finance

Build the boundary-case test set for a contested term in my field

You get a boundary-case test set for one contested term: a set of edge cases, an in-or-out verdict for every case under every rival definition you supplied, at least one case that the rival definitions actually split on, and a named residue of cases your own preferred definition leaves undecided.

You receive: A JSON boundary-case record: the contested term, a numbered case set (each case with a short phenomenon description and the feature statements it contests), a complete verdict matrix using exactly the three allowed verdict values "in", "out", and "unsettled", a four-criterion assessment (operational, essential, widelyApplicable, succinct) for each supplied definition using exactly the three allowed values "meets", "fails", and "unclear", and a residue list of the case ids your preferred definition marks "unsettled".

Part of Win Market

What's verified: STUD verifies the STRUCTURE of the boundary-case set only. It checks that the record is about the term you named; that you supplied at least your minimum number of rival definitions and that each one carries a source attribution; that you supplied at least your minimum number of cases with no case repeated; that every case is judged under every supplied definition, exactly once, with nothing left blank; that every verdict is one of exactly these three literal strings: "in", "out", "unsettled" (nothing else passes, including "yes", "no", "maybe", "+", "-", "?"); that at least your minimum number of cases are ones the rival definitions actually split on; that exactly one case is flagged as a control and that all definitions agree on it; that every other case names at least one of your supplied feature statements it contests; that every supplied definition is scored against exactly the four criteria operational, essential, widelyApplicable, and succinct, each using one of exactly these three literal strings: "meets", "fails", "unclear"; and that the residue list is exactly the set of cases your preferred definition marks "unsettled", no omissions and no extras. STUD does NOT verify that a case is well chosen, interesting, or representative of the real edges of your term. STUD does NOT verify that a verdict is CORRECT: whether a given definition truly puts a given case in or out is a reading of that definition, and readings are contested (in the source paper 52% of surveyed professionals contradicted their own stated criteria when judging cases). STUD does NOT check succinctness by word count, because the source names succinctness as a criterion but sets no numeric threshold. STUD does NOT require that every feature statement you supply be contested by some case: the source's own 20-case set left one of its 13 statements (D, on observability) untested by any example, so full coverage is not treated as a law here. STUD does NOT judge whether your preferred definition is better than its rivals. The field or discipline you name is collected as context for the operator and is not itself checked.

Opens soon

Cost20 credits
ProtectionHeld until verified delivery

This play is verified and ready. It opens soon, once sign-in and payments are live.

Example

A sample of what this play produces. Your result is generated for your inputs.

TermAI agent

Cases

Case IdC0
Is Controltrue
PhenomenonA one-page PDF price sheet listing STUD's Standard, Pro, and Ultra plans, drafted once by a person and emailed by hand to a prospective buyer, with no software step between drafting and sending.
Case IdC1
Is Controlfalse
Contests
  • A
  • C
PhenomenonSTUD's own judge worker runs a buyer's frozen acceptance checks against a submitted deliverable in a fixed, pre-written order and returns accepted or rejected, never adding, skipping, or reordering a check.
Case IdC2
Is Controlfalse
Contests
  • A
  • B
PhenomenonA coding assistant that, given a buyer's failing test suite, autonomously chooses which files to open, which edits to make, and how many times to rerun the tests, stopping on its own judgment, with no human approving individual edits.
Case IdC3
Is Controlfalse
Contests
  • A
  • C
PhenomenonA single fixed script that makes exactly one model call per request, translating one buyer-submitted paragraph and returning the text, with no second tool call and no step where the script decides anything.
Case IdC4
Is Controlfalse
Contests
  • D
PhenomenonSTUD's concierge founder-operator personally drafts a buyer's deliverable by hand for a play still gated behind manual fulfillment, then submits it to the same judge worker an agent's submission would face, and the buyer pays only if the judge accepts it.
Case IdC5
Is Controlfalse
Contests
  • A
  • B
PhenomenonA customer-support bot that retrieves matching help-center articles for an incoming question and returns the single best-matching article verbatim, choosing which article to return but never calling a second tool or taking a second step.
Case IdC6
Is Controlfalse
Contests
  • A
  • B
PhenomenonA multi-agent pipeline where a planner model breaks a buyer's request into sub-tasks, hands each sub-task to a separate specialist model, and a final model merges their outputs, with every hand-off happening without a human approving it.
Case IdC7
Is Controlfalse
Contests
  • A
  • C
PhenomenonA workflow builder lets a buyer wire together fixed, numbered steps (call model, call tool, call model again) in an order the buyer sets at build time, and the software executes exactly that fixed order for every request with no deviation.

Verdicts

Case IdDefinition IdVerdict
C0D1out
C0D2out
C0D3out
C0D4out
C1D1in
C1D2out
C1D3out
C1D4out
C2D1in
C2D2in
C2D3in
C2D4unsettled
C3D1in
C3D2out
C3D3out
C3D4out
C4D1out
C4D2out
C4D3out
C4D4in
C5D1in
C5D2unsettled
C5D3out
C5D4unsettled
C6D1in
C6D2in
C6D3in
C6D4unsettled
C7D1in
C7D2out
C7D3out
C7D4out

Definition Assessments

Definition IdOperationalEssentialWidely ApplicableSuccinct
D1meetsfailsmeetsmeets
D2unclearmeetsmeetsfails
D3meetsmeetsunclearmeets
D4meetsmeetsunclearmeets

Residue Case Ids

  • C2
  • C5
  • C6

Get early access to STUD the day it goes live.