arXiv:2402.14978
CHI 2024
4 LLM Personas

BRAINWRITE
AI-Augmented Group Ideation

Give it a challenge. Get back a complete scored idea pool — every idea rated on Relevance, Innovation, and Insightfulness by a research-validated LLM evaluation engine.

  • 4 independent LLM personas generate ideas silently — eliminating production blocking and peer judgment
  • LLM Enhancer generates conceptual blends and fills gaps the personas missed
  • Every idea scored on 3 axes (Relevance · Innovation · Insightfulness, Likert 1–5)
  • Full ranked idea pool with EXCEPTIONAL / STRONG / MODERATE / WEAK tier labels

Pipeline at a Glance

Pipeline Stages

Divergence + Convergence

2

Participant Personas

Engineer · User · Systems · Futurist

4

Ideas Per Persona

generated independently (silent phase)

6–8

LLM Enhancement

blended ideas added after human phase

6–10

Evaluation Criteria

Relevance · Innovation · Insightfulness

3

Paper Reference

Shaer, Cooper & Mokryn

CHI 2024
Two-Stage Pipeline

From Challenge to Ranked Idea Pool

Stage 1 maximises idea diversity through independent silent generation and LLM blending. Stage 2 applies a research-validated 3-axis scoring engine to every idea in the pool.

STAGE 1
Divergence
1

Problem Framer

Parses the challenge into objectives, constraints, stakeholders, and design space. Searches for domain context and analogous solved problems.

2

Silent Panel (×4)

Four personas independently generate 6–8 ideas each with no access to others' outputs — eliminating production blocking and peer judgment. Pragmatic Engineer · User Advocate · Systems Thinker · Visionary Futurist.

3

Pool Compiler

Merges all four participant idea sets into a single numbered, source-tagged shared idea pool (the virtual Brainwriting whiteboard).

4

LLM Enhancer

Reviews the full merged pool, searches for additional inspiration, and generates 6–10 new conceptually blended ideas — extending high-potential concepts and filling gaps no single participant covered.

STAGE 2
Convergence
5

Idea Evaluator

Scores every idea (participant + LLM-enhanced) on three axes: Relevance, Innovation, and Insightfulness (Likert 1–5). Calls evaluate_idea once per idea. The paper's LLM Evaluation Engine achieved Fleiss' Kappa ≥ 0.40.

6

Idea Ranker

Computes composite scores, ranks all ideas, and assigns tier labels: EXCEPTIONAL (≥4.5) · STRONG (≥3.5) · MODERATE (≥2.5) · WEAK (<2.5). Produces the TOP_IDEAS_JSON for the brief.

7

Brief Writer

Assembles the final AI-Augmented Brainwriting Report: session overview, problem profile, full ranked idea pool with scores and tier labels, evaluation summary, and source attribution.

Phase 2: Four Independent Participant Personas — The Silent Brainwriting Phase

⚙️ A — Pragmatic Engineer

Technical feasibility, existing tools, implementation steps, engineering trade-offs. Focuses on what can be built in 6–12 months.

🧑‍🤝‍🧑 B — User Advocate

Human needs, pain points, accessibility, emotional experience. Centres diverse users — including edge cases.

🔗 C — Systems Thinker

Ecosystem, dependencies, feedback loops, scalability, unintended consequences. Focuses on systemic change.

🚀 D — Visionary Futurist

Emerging trends, cross-domain analogies, radical reframing. Challenges assumptions and redefines the problem.

4 personas · 6–8 ideas each · 24–32 human ideas + 6–10 LLM-enhanced · All scored on 3-axis Likert

Sample Output

What You Get

Every idea scored and ranked — not a flat list, but a tiered pool with source attribution so you know whether each idea came from a human-like persona or from LLM conceptual blending.

BRAINWRITE — Ranked Idea Pool
Sample Output

Challenge: “Design a mental health support system for remote workers”

#1
EXCEPTIONAL
LLM-Enhanced
Score: 4.7/5

Ambient Wellbeing Sensor + AI Coach

Passive biometric sensors (HRV, typing cadence, webcam) feed a lightweight AI that detects stress patterns and surfaces micro-interventions before burnout — combining Participant A's wearable idea with Participant C's early-warning systems lens.

R: 5  ·  I: 5  ·  S: 4Relevance · Innovation · Insightfulness

#2
EXCEPTIONAL
Participant D
Score: 4.5/5

Digital "Watercooler" — Serendipitous Connection Engine

Algorithm that matches remote workers for brief spontaneous video micro-meetings based on shared interests and compatible schedules — recreating the social serendipity of office proximity that remote work eliminates.

R: 5  ·  I: 4  ·  S: 5Relevance · Innovation · Insightfulness

#3
STRONG
Participant B
Score: 4.0/5

Async Emotional Check-In Rituals

Team-configured brief async check-ins (emoji-based mood pulse + one optional sentence) woven into existing stand-up workflows — low friction, normalises emotional disclosure, creates longitudinal team mood data.

R: 4  ·  I: 4  ·  S: 4Relevance · Innovation · Insightfulness

#4
STRONG
Participant A
Score: 3.7/5

Context-Aware Focus Block Scheduler

Calendar tool that negotiates focus blocks across team calendars automatically and enforces them with "do not disturb" integrations — directly reducing cognitive context-switching as a structural stressor.

R: 4  ·  I: 4  ·  S: 3Relevance · Innovation · Insightfulness

#5
STRONG
Participant C
Score: 3.5/5

Company-Level Burnout Early Warning Dashboard

Aggregated (anonymised) signals from communication tools and calendar data feed a team-health dashboard for managers — enables systemic intervention before individual burnout becomes retention loss.

R: 4  ·  I: 3  ·  S: 4Relevance · Innovation · Insightfulness

4 personas · 28 ideas generated · 8 LLM-enhanced · 36 total scored
DIVERGE + CONVERGE
The Evaluation Engine

3-Axis Scoring Criteria

The evaluation engine from Shaer et al. (CHI 2024) scored ideas with Fleiss' Kappa ≥ 0.40 across all three criteria (29 evaluation repetitions) — moderate-to-good inter-rater agreement consistent with human expert evaluators.

Likert 1–5

Relevance

Extent the idea is connected to and appropriate for the problem statement objectives and requirements.

5

Directly addresses all core objectives, constraints respected

3

Connected to the problem but with gaps

1

Tangential or off-topic

Likert 1–5

Innovation

How original and creative the idea is; breaking away from conventional solutions or expected approaches.

5

Surprising, non-obvious, challenges existing assumptions

3

Some novelty but builds on known approaches

1

Familiar solution with little differentiation

Likert 1–5

Insightfulness

Reflects a profound and nuanced understanding of the problem statement, its root causes, and stakeholder needs.

5

Grasps latent/systemic aspects of the problem most would miss

3

Shows understanding of surface-level problem dimensions

1

Superficial; misses key problem dynamics

Human vs LLM Idea Space

  • Humans used: "people", "wearable", "screen", "work", "time", "space"
  • LLM (GPT-3) used: "users", "device", "surface", "light", "posture", "wrist"
  • Humans were more abstract and contextual; LLM more concrete and device-specific
  • LPA analysis confirmed distinct lexical profiles — why blending both is essential

Expert Alignment

  • Top GPT-4-rated ideas were generally in the top half of expert rankings
  • Expert-rejected ideas consistently appeared in GPT-4's lower tiers
  • None of the expert-discarded ideas were chosen by teams for development
  • GPT-4 evaluation complemented but did not fully replicate expert nuance
Ready to Ideate

Start a Brainwriting Session

Describe your challenge and let four thinking lenses — plus an LLM Enhancer — generate and score your idea pool in minutes.

Based on Shaer, Cooper, & Mokryn (2024) · CHI 2024 · arXiv:2402.14978