AI in a State-Regulated Evaluation Workflow

What an AI layer can and can't own in a teacher evaluation workflow.

Frontline Education · Product Designer · Collaborative Video · Professional Growth · Concept · 2025

What this is

In 2022, I shipped a collaborative video feature that let evaluators leave timestamped notes on recorded lessons, align each observation to a rubric criterion, and assign a performance rating. When they hit "Complete and Save PDF," they got a table sorted by timestamp, every observation in the order it was recorded. A teacher receiving that document had to do the synthesis work themselves.

The AI tooling to fix this didn't exist in 2022. This is the feature the export always needed.

Prototype

A working prototype built in Claude Code demonstrating the core interaction: timestamped observations on the left, AI-generated summary organized by rubric domain on the right, editable before anything goes into the teacher record.

Try the interactive prototype →

What I Did

  • Identified the gap in the original shipped feature and defined the design problem
  • Worked through what AI could and couldn't own in a high-stakes evaluation context
  • Designed the interaction and wrote the prototype brief
  • Built a working prototype in Claude Code

Methods

  • Domain analysis
  • Interaction design
  • Prototype (Claude Code)

Tools

  • Claude Code
  • Figma

The original tool

The collaborative video feature let evaluators watch a recorded lesson and build a structured record as they watched. Timestamped comments captured specific moments. A rubric panel let them align each observation to a Danielson criterion and select a performance level.

Collaborative video player showing timestamped comments panel alongside the video
The video player with timestamped comments. Each note could be aligned to a rubric criterion and rated before moving on.

The review screen collected everything into a table: timestamp, evaluator note, rubric alignment, rating. It exported to PDF in that same order.

Review screen showing timestamped observations in a table with rubric alignments and ratings
The review screen — a chronological table of every observation, organized by when it was recorded rather than by what it meant.

A teacher receiving a 12-row table sorted by timestamp had to find all the 3b observations themselves, notice when two of them pointed in different directions, and work out what the pattern meant for their practice.

The gap

The PDF reflects the order notes were taken in, not how a teacher processes feedback or how an administrator makes a summative judgment.

What both audiences need is evidence grouped by domain: all the 1a observations together, all the 3b observations together, so a pattern across multiple moments shows up in one place.

The inputs are structured. Each note has a timestamp, observation text, criterion alignment, and a rating the evaluator already selected. Reorganizing that by domain instead of timeline is a synthesis problem, and it's one AI can do well without adding anything the evaluator didn't already record.

What AI can and can't own here

Teacher evaluation in K-12 is high stakes. Evaluations inform tenure decisions, remediation plans, sometimes termination. In many states, evaluation frameworks are negotiated with unions. AI judgment in that process, even subtly introduced, would be a liability.

So I worked through where the line was.

AI can't suggest ratings mid-observation. Anchoring bias is real, and an evaluator who sees a suggested score before selecting their own isn't fully exercising their judgment. AI can't flag misaligned evidence in the moment either, which puts the tool in an adversarial position mid-task. And it can't generate a summary that introduces language or judgment the evaluator didn't already express.

What's left is reorganizing what the evaluator already collected into a form they didn't have time to produce themselves. The evaluator still writes every word that ends up in the summary.

The prototype

The interaction is a single moment: the evaluator has finished watching and aligning their notes. Left panel shows the raw observations, exactly as recorded. They click "Generate Summary." The right panel populates with a narrative organized by Danielson domain, written from their own selections.

Two-panel prototype showing timestamped observation notes on the left and AI-generated summary organized by Danielson domain on the right
Left panel: raw observations in the order they were taken. Right panel: the same evidence reorganized by rubric domain, generated on demand.

The summary doesn't flatten mixed evidence. When an evaluator recorded two 3b observations that landed differently, the summary says so. Resolving that into a tidy average would make the document less accurate, not more useful.

The right panel is editable. The evaluator reads it, revises it, owns it before anything goes into the teacher record. What gets sent is the evaluator's document.

What's next

The prototype demonstrates the interaction at the point of summary generation. The broader concept extends further: the same structured evidence that produces an individual summary could feed a district-level calibration view, where administrators can see how evaluators across a school are aligning similar evidence to different ratings over time.

That's a separate surface. The summary prototype is the piece I didn't have the tooling to build in 2022.

Key Takeaways

  • Where AI organizes versus where it judges is a decision you make, not a default. Letting AI suggest ratings or flag misalignment mid-task puts it in an adversarial position with the evaluator, in a context where the output has legal and institutional consequences. This design draws the line at synthesis. The evaluator's judgment is the only judgment in the document.
  • Output format matters as much as the data underneath it. The original tool collected accurate evidence in an order nobody could use. Reorganizing by domain instead of timestamp is a small structural change that changes how the document gets read.
  • Sometimes the right move is to leave the ambiguity in. When evidence pointed in different directions, the summary names that instead of averaging it away. A document that shows the real complexity of what was observed holds up better under scrutiny than one that smooths it out.