Give it a challenge. Get back a complete scored idea pool — every idea rated on Relevance, Innovation, and Insightfulness by a research-validated LLM evaluation engine.
Pipeline Stages
Divergence + Convergence
Participant Personas
Engineer · User · Systems · Futurist
Ideas Per Persona
generated independently (silent phase)
LLM Enhancement
blended ideas added after human phase
Evaluation Criteria
Relevance · Innovation · Insightfulness
Paper Reference
Shaer, Cooper & Mokryn
Stage 1 maximises idea diversity through independent silent generation and LLM blending. Stage 2 applies a research-validated 3-axis scoring engine to every idea in the pool.
Parses the challenge into objectives, constraints, stakeholders, and design space. Searches for domain context and analogous solved problems.
Four personas independently generate 6–8 ideas each with no access to others' outputs — eliminating production blocking and peer judgment. Pragmatic Engineer · User Advocate · Systems Thinker · Visionary Futurist.
Merges all four participant idea sets into a single numbered, source-tagged shared idea pool (the virtual Brainwriting whiteboard).
Reviews the full merged pool, searches for additional inspiration, and generates 6–10 new conceptually blended ideas — extending high-potential concepts and filling gaps no single participant covered.
Scores every idea (participant + LLM-enhanced) on three axes: Relevance, Innovation, and Insightfulness (Likert 1–5). Calls evaluate_idea once per idea. The paper's LLM Evaluation Engine achieved Fleiss' Kappa ≥ 0.40.
Computes composite scores, ranks all ideas, and assigns tier labels: EXCEPTIONAL (≥4.5) · STRONG (≥3.5) · MODERATE (≥2.5) · WEAK (<2.5). Produces the TOP_IDEAS_JSON for the brief.
Assembles the final AI-Augmented Brainwriting Report: session overview, problem profile, full ranked idea pool with scores and tier labels, evaluation summary, and source attribution.
⚙️ A — Pragmatic Engineer
Technical feasibility, existing tools, implementation steps, engineering trade-offs. Focuses on what can be built in 6–12 months.
🧑🤝🧑 B — User Advocate
Human needs, pain points, accessibility, emotional experience. Centres diverse users — including edge cases.
🔗 C — Systems Thinker
Ecosystem, dependencies, feedback loops, scalability, unintended consequences. Focuses on systemic change.
🚀 D — Visionary Futurist
Emerging trends, cross-domain analogies, radical reframing. Challenges assumptions and redefines the problem.
4 personas · 6–8 ideas each · 24–32 human ideas + 6–10 LLM-enhanced · All scored on 3-axis Likert
Every idea scored and ranked — not a flat list, but a tiered pool with source attribution so you know whether each idea came from a human-like persona or from LLM conceptual blending.
Challenge: “Design a mental health support system for remote workers”
Ambient Wellbeing Sensor + AI Coach
Passive biometric sensors (HRV, typing cadence, webcam) feed a lightweight AI that detects stress patterns and surfaces micro-interventions before burnout — combining Participant A's wearable idea with Participant C's early-warning systems lens.
R: 5 · I: 5 · S: 4│Relevance · Innovation · Insightfulness
Digital "Watercooler" — Serendipitous Connection Engine
Algorithm that matches remote workers for brief spontaneous video micro-meetings based on shared interests and compatible schedules — recreating the social serendipity of office proximity that remote work eliminates.
R: 5 · I: 4 · S: 5│Relevance · Innovation · Insightfulness
Async Emotional Check-In Rituals
Team-configured brief async check-ins (emoji-based mood pulse + one optional sentence) woven into existing stand-up workflows — low friction, normalises emotional disclosure, creates longitudinal team mood data.
R: 4 · I: 4 · S: 4│Relevance · Innovation · Insightfulness
Context-Aware Focus Block Scheduler
Calendar tool that negotiates focus blocks across team calendars automatically and enforces them with "do not disturb" integrations — directly reducing cognitive context-switching as a structural stressor.
R: 4 · I: 4 · S: 3│Relevance · Innovation · Insightfulness
Company-Level Burnout Early Warning Dashboard
Aggregated (anonymised) signals from communication tools and calendar data feed a team-health dashboard for managers — enables systemic intervention before individual burnout becomes retention loss.
R: 4 · I: 3 · S: 4│Relevance · Innovation · Insightfulness
The evaluation engine from Shaer et al. (CHI 2024) scored ideas with Fleiss' Kappa ≥ 0.40 across all three criteria (29 evaluation repetitions) — moderate-to-good inter-rater agreement consistent with human expert evaluators.
“Extent the idea is connected to and appropriate for the problem statement objectives and requirements.”
Directly addresses all core objectives, constraints respected
Connected to the problem but with gaps
Tangential or off-topic
“How original and creative the idea is; breaking away from conventional solutions or expected approaches.”
Surprising, non-obvious, challenges existing assumptions
Some novelty but builds on known approaches
Familiar solution with little differentiation
“Reflects a profound and nuanced understanding of the problem statement, its root causes, and stakeholder needs.”
Grasps latent/systemic aspects of the problem most would miss
Shows understanding of surface-level problem dimensions
Superficial; misses key problem dynamics
Describe your challenge and let four thinking lenses — plus an LLM Enhancer — generate and score your idea pool in minutes.
Based on Shaer, Cooper, & Mokryn (2024) · CHI 2024 · arXiv:2402.14978