Joke Rater.

How scoring works

A professional-comedian AI reads your bit and grades its craft — the writing and structure — across the dimensions below. It then rolls the structural and content dimensions into one number: the Craft / Structure Score.

This is a craft-and-structure score, not a prediction of how funny the bit is or how a room will react. Comedy is audience- and moment-dependent — use this to sharpen the writing, not to forecast laughs.

The composite

Only the structural and content dimensions feed the composite (weighted — core structure counts more than speculative "potential" dimensions). The audience / context dimensions are reported next to it, never folded in — a risky or niche bit shouldn't drag down a craft score.

  • 0 –39 Weak
  • 40 –59 Developing
  • 60 –79 Solid
  • 80 –100 Strong

Structural / mechanical

Setup

weight 3

Clarity and effectiveness of the premise-setting line(s).

The setup establishes one clear premise and the assumption the audience will make. A muddy or overloaded setup steals energy from the punchline; good setups are economical and land on the word that frames the expectation.

0–1
Premise is unclear or carries several competing ideas.
2–3
Premise is readable but takes a detour or an extra beat to land.
4–5
One clean premise; the assumption to be broken is unmistakable.

Build-up

weight 2

How well the bit escalates or develops before the punchline.

Between premise and punchline the bit should raise stakes or sharpen the picture, each sentence earning the next. Flat middles let tension leak out before the laugh.

0–1
No development — premise sits still until the punchline.
2–3
Some development, but a beat or two doesn't add pressure.
4–5
Every line tightens the picture or raises the stakes.

Punchline

weight 3

Strength of the punchline itself.

The punchline is the reveal that breaks the assumption the setup planted. Strength comes from a clean, surprising connection — not from volume, shock alone, or explaining the joke.

0–1
No clear reveal, or it just restates the setup.
2–3
There is a reveal, but the connection is loose or over-explained.
4–5
Sharp, clean break from the expectation; nothing wasted.

Punchline placement

weight 2

Whether the key/funniest word lands at the very end of the sentence, not buried mid-line.

Standard stand-up writing puts the funniest word — the reveal — as the last word of the sentence. Words after it bury the laugh and signal the audience to wait.

0–1
Key word is mid-sentence with a trail of words after it.
2–3
Reveal is near the end but a few words still follow it.
4–5
The sentence ends on the reveal word.

Misdirection strength

weight 3

How effectively the setup creates a false expectation the punchline subverts (vs. just "fact + tag").

A strong joke makes the audience confidently expect one thing, then the punchline reveals another reading that was there all along. 'Fact + tag' — a true observation plus a quip — is lighter than genuine misdirection.

0–1
No false expectation set; it's an observation with a comment.
2–3
A mild expectation is set but the turn is soft.
4–5
The setup commits the audience one way; the reveal flips it cleanly.

Economy

weight 2

Word efficiency in the setup; flags unnecessary words delaying the punchline.

Every word in the setup that isn't doing work delays the punchline and dilutes it. Cut qualifiers, throat-clearing, and redundant scene-setting.

0–1
Noticeable filler, hedging, or repeated scene-setting.
2–3
Mostly tight, with a couple of trimmable words.
4–5
Nothing removable without losing sense or the picture.

Tag potential

weight 1

Does the bit leave room for a second/third laugh immediately after the main punchline?

A tag is a second punchline that runs off the same setup, landing right after the first with no new setup. Bits that leave an angle unspent have tag potential.

0–1
Premise is fully spent; no obvious second angle.
2–3
A possible tag exists but would need a fresh setup.
4–5
One or more clear tags sit right off the existing setup.

Callback potential

weight 1

Does the bit plant something (phrase, image, premise) reusable later in a longer set?

A callback re-uses an earlier phrase, image, or premise later in the set, earning a laugh from recognition. Vivid, nameable elements travel well; generic ones don't.

0–1
Nothing distinctive enough to bring back.
2–3
A phrase or image could be reused with some setup.
4–5
A vivid, nameable element is begging to return later.

Rhythm / cadence

weight 2

Sentence rhythm; patterns like the rule of three; spoken-rhythm vs. prose-like phrasing.

Spoken comic rhythm favors short, front-loaded sentences and patterns like the rule of three — two to establish, one to break. Prose-like, comma-heavy phrasing reads fine but doesn't land out loud.

0–1
Long, comma-heavy sentences; reads like an essay.
2–3
Speakable, but the stresses don't fall where the laughs are.
4–5
Clear spoken cadence; beats land on the funny words.

Content / craft quality

Specificity

weight 2

Vivid, specific concrete detail vs. vague general observation.

Concrete, specific detail is funnier and more credible than general observation. Specificity signals a real point of view.

0–1
General and abstract; could be about anything.
2–3
Some concrete detail, some vague patches.
4–5
Precise, sensory detail that makes the picture undeniable.

Surprise / predictability

weight 3

How unpredictable the punchline is relative to where the setup points — distinct from misdirection strength.

Measures how far the punchline sits from where the setup pointed. A joke can technically subvert expectation yet still feel predictable if the subversion is a familiar trope; genuine surprise is a turn the audience didn't have on their list.

0–1
The punchline is the first place most people would go.
2–3
Not obvious, but the type of turn is a known trope.
4–5
A turn the audience wouldn't have listed, yet it fits.

Escalation potential

weight 1

Whether the premise supports building across multiple beats vs. being a single-hit line.

Some premises are single-hit lines; others can be pushed through escalating beats, each more absurd or higher-stakes than the last. This rates whether the premise has somewhere to go.

0–1
One-and-done; the idea is fully used up.
2–3
Could take one more beat before it runs out.
4–5
The premise clearly supports a rising run of beats.

Persona / voice consistency

weight 2

Does the tone match a coherent comedic point of view, or could this joke be anyone's?

Strong material sounds like it could only come from one specific comic — a consistent attitude, values, and voice. 'Could be anyone's joke' is a note, not a compliment.

0–1
Generic voice; no discernible point of view.
2–3
A point of view is visible but wavers in tone.
4–5
Unmistakably one voice and attitude throughout.

Audience / context

Reported for calibration; not part of the composite.

References & vocabulary

context

Identify the references and vocabulary level the bit uses.

Names the references (people, events, brands, subcultures) and the vocabulary level the bit leans on, so the writer can see what knowledge they're assuming. Not a quality judgement — a map.

0–1
No outside references; everyday vocabulary.
2–3
A few mainstream references or some elevated vocabulary.
4–5
Several references or specialised vocabulary that carry the bit.

Accessibility (self-containment)

context

How much shared context a listener needs to get the joke — self-contained vs. requires niche knowledge.

Rates only how self-contained the joke is: can a general listener follow it with no outside knowledge, or does it need a specific reference, jargon, or niche context? Higher = more self-contained. This never profiles audiences by group.

0–1
Needs specific niche knowledge to land at all.
2–3
Works for most, but one reference narrows the room.
4–5
Fully self-contained; no outside knowledge required.

Accessibility here measures only how self-contained the joke is — how much outside knowledge a listener needs to follow it. It never scores or profiles audiences by demographic group.

Restriction

context

Flags content/language that may need audience-appropriateness consideration; maps to an age-rating band.

Flags language or content — profanity, sexual content, slur-adjacent edge, graphic violence — that affects where the bit can be performed or published, and maps to an age-rating band (general / teen / mature).

0–1
Clean; performable anywhere (general).
2–3
Some profanity or adult themes (teen).
4–5
Strong language or explicit content (mature).

Topicality vs. evergreen

context

Dependent on a current event/trend (dates quickly) vs. timeless material.

Rates how tied the bit is to a current event or trend. Topical material is sharp now but dates fast; evergreen material keeps working for years.

0–1
Evergreen; no time-sensitive hook.
2–3
Lightly tied to a current theme; a few months of shelf life.
4–5
Depends on a specific current moment; dates within weeks.

Edginess / risk

context

How far the material pushes into taboo/provocative territory — for the comic's calibration and for moderation.

Rates how far the material pushes into taboo or provocative territory — useful for the comic's own calibration and for platform moderation. High edginess isn't bad, but it should be a deliberate choice.

0–1
No taboo content; broadly inoffensive.
2–3
Touches a sensitive area but doesn't dwell there.
4–5
Leans hard into provocative or taboo territory.

Physical / act-out potential

context

Whether phrasing suggests a moment that would benefit from a physical beat or voice shift.

Flags moments where the phrasing implies a physical beat or a voice/character shift a performer could act out. A forward hook toward delivery coaching; text can only suggest it.

0–1
Purely verbal; nothing implies physicality.
2–3
One spot could carry a gesture or voice change.
4–5
The bit clearly wants an act-out or character voice.

The 0–5 anchor descriptions are informed by stand-up craft literature and are still being validated with working comedians. Treat the numbers as a structured second read, not a verdict.