Skip to content
NexCell Stage Two / reasoning dossier

The eleven answers.

Eleven prompts on how you would design and build a dashboard for watching live AI conversations. Answered below, and linked to where each one runs.

note

The brief asked for reasoning, not code. This document answers all eleven prompts. It also builds the dashboard those answers describe, so the reasoning can be watched running rather than taken on faith. Each answer carries a link into the exact place it is demonstrated.

Everything runs on a client-side simulation by default. Bring your own Groq or OpenRouter key for real streaming AI, browser to provider direct, the key in localStorage and never on a server of ours. No paid services, no server dependency. Built for the NexCell Stage Two assessment.

01prompt 01

Component structure.

Three panes under one command bar, mission-control layout. Each pane is a container that self-drives from a hook; the page only mounts them.

  • ConversationList (left): a dense, virtualized table of every live thread. Owns selection only. Reads the filtered, triage-sorted list and writes selectedId.
  • ConversationTranscript (center): the glass box. Renders one thread, owns scroll anchoring and the streaming view.
  • OperatorPanel (right): AIStatusIndicator (the honest state machine) plus OperatorControls (take over, release, composer, canned replies).
  • CommandBar (top): search, status filters, live metrics, connection state, theme, and the demo controls.

Rows and bubbles are memoized leaf components, so a frame that touches one thread does not re-render the rest.

See it livethe labelled system diagram
02prompt 02

State and data flow.

Three stores, one rule: server data has exactly one home.

  • Fetched and subscribed data lives in TanStack Query, keyed ["conversations"] and ["conversation", id]. The socket keeps no copy of its own; incoming frames are written straight into the Query cache with setQueryData.
  • Ephemeral UI state lives in Zustand: selectedId, status filter, search, connection state, live-mode config. It stores whether a key exists, never the key bytes.
  • Shareable state lives in the URL (selectedId, filter, search), so a thread is a link.

Server state is never mirrored into Zustand. One socket, multiplexed by subscription, feeds one cache.

See it livethe store and socket wiring
03prompt 03

Figma to a responsive interface.

Tokens first. Palette, type scale, and spacing become CSS custom properties, so a Figma value maps to one variable rather than a scatter of literals. Desktop is an asymmetric 26/50/24 three-pane grid. Below the breakpoint the panes do not shrink into each other; they become routed views (the list, then tap through to a transcript), because a forty-character-wide transcript is useless.

Cross-breakpoint check: drag the viewport continuously from 1440 to 320. Nothing overflows the x-axis because each pane scrolls inside its own container and the body clips horizontally; the transcript reflows and the table drops low-value columns instead of truncating rows.

See it liveresize the window from 1440 down to 320
04prompt 04

API contract: what the frontend needs, and one response shape.

The list needs a scannable summary with no message bodies; the detail needs the thread plus the AI live state. The field everything hinges on is seq: a monotonic-per-conversation counter that orders frames, dedups overlapping REST and socket delivery, and detects gaps after a reconnect.

GET /conversations?status=needs_human&limit=50

{
  "items": [
    {
      "id": "cv_0007",
      "code": "DP-0007",
      "status": "needs_human",
      "member": { "handle": "+44 •••• 1183", "channel": "sms" },
      "scenario": "parcel dispute",
      "aiPhase": "waiting",
      "risk": "high",
      "waitingMs": 41200,
      "lastSeq": 87,
      "unreadCount": 2,
      "lastMessage": { "preview": "I still have not had a refund", "role": "member", "timestamp": 1720001234000 }
    }
  ],
  "nextCursor": "cv_0007"
}

risk and waitingMs sit in the row because the triage sort runs client-side; lastSeq lets the socket resume this row on reconnect without a full refetch.

See it livethe contract panel
05prompt 05

Endpoints: what each returns and one edge case.

  • GET /conversations returns { items, nextCursor }. Edge: limit is capped 1 to 100 (422 over), status is validated against the enum (422 on junk).
  • GET /conversations/{id} returns the detail plus the newest N messages. Edge: 404 for missing versus 403 for another tenant, with the deliberate note that a 403 leaks existence.
  • GET /conversations/{id}/messages returns { items, hasMore, nextCursor }. Edge: after=lastSeq is the reconnect catch-up; if more than a page is waiting it returns hasMore, never a silent truncation.
  • POST /conversations/{id}/messages returns 202 and the echoed message with its seq. Edge: an Idempotency-Key header is required (missing is 422); a double-submit with the same key posts once.
  • POST /conversations/{id}/takeover is compare-and-set. Edge: the operator who loses a two-operator race gets 409.
See it livethe endpoint table (a real FastAPI backend implements it)
06prompt 06

Loading and error states.

Every pane owns its loading, empty, and error state, scoped so one failure does not blank the screen. Skeletons match the real row and bubble geometry, so nothing reflows on arrival. Empty is a designed prompt, not a void. A dropped socket surfaces as a reconnecting banner.

Why it matters: in a live tool a dead feed and a quiet conversation look identical, and that ambiguity is how a waiting customer gets ignored. To make every state reviewable, the command bar has PROVE IT toggles that force any pane into loading, empty, or error on demand.

See it livethe PROVE IT force-state toggles in the command bar
07prompt 07

Performance: hundreds of conversations, thousands of messages.

  • Virtualize both lists with @tanstack/react-virtual. The table holds roughly twenty visible rows in the DOM out of hundreds; the transcript measures variable-height bubbles dynamically and renders only what is on screen. Rows and bubbles are memoized.
  • Batch the socket. Inbound frames are coalesced in a buffer and flushed inside a single requestAnimationFrame, so commits are capped at one paint per frame regardless of inbound rate. A 200-frame burst becomes one commit, not 200 re-renders.
See it livethe virtualized list, then hit BURST
08prompt 08

Testing plan: 3-4 high-signal checks.

All four checks run today. A 30-test Vitest suite covers the engine and data layers (pnpm test), React Testing Library covers the render states, and a Playwright smoke run covers the flows (pnpm e2e). None of this is a promise.

  • Engine invariants (shipped, Vitest): a suite run with pnpm test asserts the properties the whole live layer rests on: seq stays contiguous per conversation, an operator send is idempotent under a repeated idempotency key, takeover is a compare-and-set with a single winner, and empty text is rejected. This is the logic most likely to break and least visible when it does, so it is the part actually under test now.
  • Render states (shipped, RTL): React Testing Library asserts the transcript, empty states, status glyphs and skeletons render correctly. One guards the honest-AI rule directly: an AI message with no calibrated confidence renders no percentage, and a real score renders it.
  • Live update and reconnect (shipped, integration): a test drives the real conversation source and asserts an operator send delivers exactly one frame at seq lastSeq plus one, a repeated idempotency key is deduped, and a dropped socket produces a genuine seq gap on reconnect that the client backfills. MSW-style mocking of the production WebSocket is the planned next step, since the deployed demo simulates that transport.
  • Responsive (shipped, Playwright): a smoke run loads the landing and dashboard with zero console errors, follows a deep-link into a conversation, and asserts no horizontal overflow at 375px.
See it livethe Vitest suite and reproduce steps
09prompt 09

Accessibility and trust.

Accessibility: the transcript is announced through a hidden aria-live log that speaks completed messages only, never per streaming token; the visible virtualized viewport is deliberately not a live region, because row recycling on scroll would spam a screen reader. Takeover moves focus to the composer, every icon control is labelled, and reduced motion is honored.

Trust: the AI status is bound to real events, not a decorative spinner. Thinking means real work is in progress and names the action; replying means tokens are actually streaming; stalled means a dwell threshold was crossed and a human is asked to step in. We show provenance instead of an invented confidence number, and we distinguish AI drafted from delivered to member. The product is named for exactly this: the opposite of a black box.

See it livethe AI status panel on the right
10prompt 10

Risks and trade-offs.

  • Update flood. A busy queue can push more frames than React can paint. Mitigation: rAF batching plus latest-wins status with a minimum dwell so rapid flips do not strobe. Trade-off: sub-frame intermediate states are dropped, which for a monitoring view is the right call.
  • Optimistic divergence. An optimistic send can disagree with the server. Mitigation: an idempotency key and reconcile-on-echo; a failure flips the message to failed rather than lying about success.
  • Reconnect data loss. A dropped socket can miss frames. Mitigation: persist the last seq per subscription, resume with after=seq, and REST-backfill on a gap, deduping by seq. The drop is surfaced, never hidden.
See it liveDROP SOCKET, then race a takeover
11prompt 11

Real-time AI UI: smooth, honest, never jumping or falling behind.

This is the whole live layer, not a feature bolted on.

  • Smooth: coalesce frames per conversation and flush inside one requestAnimationFrame; status transitions are latest-wins with a minimum dwell so they do not strobe.
  • No jump: stick-to-bottom via an IntersectionObserver sentinel. Pinned to the bottom, new text autoscrolls; scrolled up, new messages raise a "N new" pill instead of yanking the view. Prepended history is anchored by measuring scroll height before and after.
  • Never behind: text and status arrive as separate frames and can reorder under load, so both are reconciled by seq; on reconnect, resume after=seq and backfill any gap.
  • Honest: status reflects real model events, with a stall timeout that flags a hung agent so a human steps in.
See it livehit BURST to flood, then DROP SOCKET to sever and resume

Stack. Next.js App Router, React 19, TanStack Query (cold store) and Zustand (hot state), @tanstack/react-virtual for both lists, a swappable ConversationSource with two impls (simulation and bring-your-own-key live AI), and a reference FastAPI backend that implements the contract end to end.

Reproduce. pnpm install, then pnpm dev, then open /dashboard. The backend is runnable locally under backend/. Full notes on the build and the honest constraints are on the guide.