Daily support for the people who absorb everyone else’s stress

A conversational agent for first responders, healthcare workers, caregivers, and teachers. Jobs like these take a mental toll, and typically don’t offer a way to recognize or deal with it.

This work was done inside a $100M stealth healthcare startup founded by former UnitedHealth Group leaders, where I drove the design thinking, guided the team’s direction, and influenced the product’s priorities.

The product does three things:

  • ABuild self-awareness — help people see what is actually going on with them
  • BDevelop personal agency — turn that into better choices about health and life
  • CLower claims costs — cut what the employer spends on insurance

What I Did

Product definition for an AI product no one had defined.

Outcomes

  • 1

    A New Team

    Built the system that reads how a person is actually doing: seven behavioral signal types picked up from live conversation, and replies written from that state, not just their words. Two ML notebooks defined it in code, and leadership formed a five-person conversational dynamics team from them.

  • 2

    60% Less Drop-Off

    Projected 60 percent less early drop-off, from a four-tier progression that matches each person’s self-awareness and stays or scales with their pace.

  • 3

    The 90-Day Plan

    Mapped seven disconnected services into one picture of the system, showing that every engineer was independently building the same unnamed thing, an agent. The map became the 90-day plan.

  • 4

    A Memory Team

    Championed user memory as core to the product, so the tool meets someone where they left off instead of starting over every session, the work that drove leadership to form a dedicated memory team.

  • 5

    A Defined Scope

    Designed where the product stops: the scope of what it helps, what it hands to human care, and what it holds back in a reply. Restraint built as a feature, not a filter.

I joined a team of machine learning engineers who had a working chatbot and no product practice. They could make it faster and cheaper. No one could say what it should do, what it should never do, or whether it was any good. That question sat unanswered in every meeting, including the ones with the consulting therapist hired to answer it.

My work was product definition. I wrote down what the product was for, where it stopped, and how it should sound. I turned the therapist’s expertise into rules the team could build against. And I made quality checkable, so “is it good” became a question with an answer instead of an argument.

Everything below is that work, rebuilt end to end as a working system.

Test bench: personas, scenarios and governance rules on the left, the conversation in the middle, and the signals and decisions behind each reply on the right

Conversational AI Test Bench

Watch the AI think. Pick a person, pick a scenario, turn signals on or off, and send. Every reply comes back with its reasoning attached, so you can see what the model read, what it decided, and what it held back. Most AI behavior gets decided in code nobody can see. I built the bench to put it in the open.

Open the Test Bench (opens in a new tab)
Therapist review console: past sessions listed on the left, the conversation in the middle with an expandable reasoning trace, and a review panel on the right for scoring, notes and flagging

Human Review Console

Real therapists shape this model. They read whole conversations, grade them, and flag and discuss the ones that worry them. Their judgment becomes the standard the machine gets measured against. That’s how you stay ahead of AI slop: have the experts work inside the loop.

Open the Review Console (opens in a new tab)
Evaluation harness: an evaluation terminal on the left and a scorecard comparing two models on quality, cost and speed, with panels below for the scenario set, the models tested and a grading run

Pipeline Evaluation

I wanted to know what quality actually costs, so I ran the same 45 scenarios through an expensive model and a cheap one, graded every reply against the product’s own rules, and put a price on every conversation. The result surprised me.

Read the Case Study

I documented the whole build as it happened, includes decisions and the dead ends (which I call “where the road hasn’t been built yet”). Sixteen sections, starting below. Start reading

HealthBot Case Study


My work ran across the whole agentic system: who it was for, how it talked, how it stayed therapeutic, how it was built, and where it had to stop.

What follows is a mix: my thinking into the artifacts that drove the work, and the outcomes that shipped out of them.

View HealthBot Test Bench

Have a live conversation with the agent here, or use the full test bench to see all the decisions made by the agent.

The Unmet Need,
and How It Works


Premise

Daily support for people whose jobs create accumulated stress: first responders, healthcare workers, government employees, caregivers.

Fifteen minutes a day of focused conversation, sustained over months, changes how people handle stress.

Business model

Unaddressed stress surfaces as burnout, lost quality of care, and turnover, and it lands on the employer’s . The bigger the issue, the bigger the cost, especially once chronic conditions take hold.

Clients pay nothing until their total cost of care drops, measured against employees who don’t use the product.

Therapeutic model

: future-oriented, goal-focused, evidence-based. I worked alongside the company’s consulting therapists to keep the conversation inside that discipline.

HealthBot draft UI — chat home screen
Can't sleep? Give HealthBot a try.

Who This Helps,
Who It Doesn’t


This section defines the product: who it helps, who it hands to human care, and the markets it fits in.

Where it helps

  • Burnout, compassion fatigue
  • Caregiver strain
  • Questions about who you are after a change
  • Restarting goals after a setback
  • Self-doubt, isolation
  • Stress-related sleep loss
  • Talking through a hard shift or call
  • Workplace pressure, role changes

Where it might help

  • ADHD
  • Anxiety without a diagnosis
  • Chronic pain or illness
  • Eating disorders (mild)
  • Grief and loss
  • Low mood that isn’t clinical depression
  • Mild OCD
  • Reducing habits (drinking, scrolling, spending)
  • Relationship strain, separation, breakups

Where it hands off

  • Active suicidality
  • Anything needing diagnosis or prescription
  • Bipolar disorder
  • Clinical depression
  • Medical issues masquerading as mental ones (thyroid, sleep apnea, diabetic crises)
  • Schizophrenia, psychosis, dissociation
  • Substance use disorders
  • Trauma needing licensed care
Comparable Products and Services A map of everything solving nearby problems.It makes direction discussable: what we are, what wearen’t, what we’re becoming. Skip the exercise, and youdrift into becoming one of these without ever deciding to. Behavioral Change Crisis Support Mindfulness + Meditation Professional Therapy Self-Reflection HealthBot Current / Future Action/Change Reflection/Understanding Professional Interaction Self-Guided Goal Orientation Guidance Suicide Prevention Hotline Mental Health Crisis Helpline Text-Based Crisis Support Individual Therapy Group Therapy Online Therapy Platforms Goal-Setting Worksheets Habit Trackers Self-Monitoring Logs Guided Journal Prompts Mood Tracker Apps Gratitude Journals Guided Meditation Apps Relaxation or Sleep Apps Breathing Exercises HealthBot’s Current State HealthBot’s Future State
Comparable Products and Services
Value of Solution-Focused Brief Therapy (SFBT) Over Time The long-term compounding benefits of repeated self-reflection.SFBT works on the present: daily management and awareness of moods, not probing childhood or past trauma. More Time Spent No Time Spent Minor Life-Threatening Self-Reflection Time Health Condition Severity Future Benefits: Minor Immediate Benefits: Compound Over Time and Effort Future Benefits: Life-Threatening Future Benefits (Minor) Immediate Benefits Future Benefits (Life-Threatening)
Value of Solution-Focused Brief Therapy

What to Build,
and Where It Can Grow


A good strategy continually asks and re-asks these questions:

What does it cost to run?

Running a conversation costs pennies, even counting the therapists who review the model’s responses. Therapy sessions run well over a hundred dollars, and are out of reach or out of mind for many frontline workers.

The product doesn’t replace therapy; it reaches the people who were never going to book it, and at a price that works across a workforce.

How do we prove it works?

The early signals have to carry the proof until the claims data can: daily use, direct feedback, conversations getting more honest. If those hold, the health care spending drop follows.

Will anyone trust it?

Trust has to be structural and visible. Some workforces require auto-deletion or encryption, so a vulnerable conversation can never become evidence in an incident. And aggregate learning needs names removed, groups too large to identify anyone.

US Healthcare Spending, 2010 to 2033 Actual spending through 2024 and federal projections after it. The shaded band is the 90 percent of spending that goesto people with chronic conditions, the population this product is built to reach earlier. $2T $4T $6T $8T 2010 2015 2020 2024 2029 2033 Trillions of dollars per year $5.3T in 2024 $15,474 per person $8.6T by 2033 $24,200 per person 90% of spending goes to people with chronic conditions projected actual $7.7T with a 10% reduction about $860 billion saved that year Actual (CMS) Projected (CMS) With a 10% reduction Spending on people with chronic conditions The Promise of Change: Reduce that spending 10 percent and the country saves about $860 billion a year by 2033.
Rising Cost of Healthcare
Science and Scalability Rigorous care costs too much and requires a referral, and the free tools have little evidence behind them.HealthBot is built to be both: backed by evidence, and available to everyone. Broad Access Selective Access Low-Rigor High-Rigor Accessibility Science-Backed Ongoing In-Person Therapy Group Therapy Online Therapy Platforms EAP Counseling Sessions 988 Suicide and Crisis Lifeline Crisis Text Line Peer Support Programs Expressive Writing Gratitude Journaling Mood Tracker Apps Free Chatbots MBSR 8-Week Course Guided Meditation Apps Breathing Exercise Apps Clinician-Delivered CBT Digital CBT Programs Habit Trackers Lifestyle Apps HealthBot:Broad Access and High Rigor Behavioral Change Crisis Support Mindfulness + Meditation Professional Therapy Self-Reflection
Scalability of Solutions
Benefit and Risk by Feature* Every feature is weighed for benefit and risk, then decided: do, manage, hold, or avoid. *Illustrative: an example set of features, not the whole backlog. High Benefit Low Benefit Low Risk High Risk Benefit Risk Do / Prioritize high value, safely held Do, But Manage Risk worth it, with safeguards funded Low Priority safe, and forgettable Avoid / Investigate risk without the payoff Daily Check-Ins Memory with User-Visible Controls Scaffolded Openings by Readiness Deep Reframing Conversations Population Distress Mapping Working near the Clinical Boundary Generic Wellness Content Streaks and Badges Diagnosis-Shaped Features Unsolicited Belief Content Handling Crisis instead of Routing It FeatureBenefitRiskCall Daily Check-InsHighLowDo Memory, User-ControlledHighLowDo Scaffolded OpeningsHighLowDo Deep ReframingHighHighManage Population MappingHighHighManage Clinical Boundary WorkMedHighManage Generic Wellness ContentLowLowHold Streaks and BadgesLowLowHold Diagnosis-Shaped FeaturesLowHighAvoid Unsolicited Belief ContentLowHighAvoid Crisis HandlingLowHighAvoid Benefit and Risk Reasoning
Benefit and Risk by Feature
Population Distress Mapping: Across a Workforce* The same classifiers that read one conversation can read a workforce. Emotional state observed rather than surveyed,compared across companies, industries, roles, and regions, to see where support is working and where to aim next. *Illustrative data. The vision: enough conversations, honestly classified, become a live picture of how a workforce is actually doing:which roles are carrying the most, where support is working, and where to aim it next. High distress (pain, grief) Low emotion (calm, stable) Low reflection (reactive, unaware) High reflection (introspective, growth-oriented) Emotional Distress Self-Reflection Intensity Raw Pain suffering Processing Pain distressed, but working Avoidance suppression, disengagement Emotional Clarity resolving, accepting, reframing Claims Processors ER Nurses Firefighters ICU Staff K-12 Teachers 911 Dispatchers
Population Distress: Workforce
Population Distress Mapping: Across a Nation* Zoomed out to a national population, the clusters stop being job titles. Geography, life stage, shift patterns, economic events:hotspots along dimensions no survey would think to cut, found by the model rather than predefined by people. *Illustrative data. The vision at scale: across a large enough population, the same classification starts finding what nobody thought to look for:patterns that could drive new products, new assistance, and new understanding of how people are really doing. High distress (pain, grief) Low emotion (calm, stable) Low reflection (reactive, unaware) High reflection (introspective, growth-oriented) Emotional Distress Self-Reflection Intensity Raw Pain suffering Processing Pain distressed, but working Avoidance suppression, disengagement Emotional Clarity resolving, accepting, reframing Adults 18 to 25 Family caregivers New parents Night-shift workers Recently unemployed Rural counties
Population Distress: Nation

HealthBot’s product-level strategy, based on Google’s

Frontline Workers and Government Employees


Reconstructed from real clients and the municipalities the product served, then used as seeds for synthetic personas and simulated conversations — the foundation for sample dialogues, scenario generation, and mapping how conversations actually go.

Jakob B.

Firefighter + Hazmat Technician
Missoula, MT

High-Risk RoleRecoveringStoic

Jakob B.
  • Recovering from second-degree burns taken during a chemical plant emergency
  • Talks about getting back to light duty, not about the doubts that come with the recovery
  • Thinks admitting it is weakness; goes quiet instead
  • He'll use it if it doesn't feel like therapy

Primary pattern

Bravado as armor. Says “I'm good” even when he isn't.

Evelyn R.

DMV Document Examiner + Title Clerk
San Diego, CA

CaregiverKids and Aging ParentsWorking Parent

Evelyn R.
  • Primary caregiver for two parents with advanced dementia
  • Mother of two teenage sons; the home shift starts the moment the work shift ends
  • Didn't know she was burned out; a friend told her
  • Thinks she's failing, not that she's running on empty

Primary pattern

Self-erasure. Cuts out her own needs first.

Kyler Z.

Advanced EMT
Cobb County, GA

High-AdrenalineJob Is IdentitySkeptical

Kyler Z.
  • Good at the job and addicted to the urgency of it
  • Coworkers worry he's burning himself out; he'd tell you he's fine
  • Told he's burned out, he'd argue the point or call it a bad shift
  • Sees calm as boredom; he levels out on the next call
  • He'll use it if it shows him numbers

Primary pattern

Deflection through busyness. Action instead of reflection.

Personas Give One Dimension.
The Model Sees Thousands.


I’ve used personas for years. They hold demographics, scenarios, and attitudes, and they help a team think about a broad set of users. The archetype model (see right) taught me a second way of thinking: instead of who someone is, read what state they’re in, their mindset and their proficiency, and adapt the product to it.

Working with an LLM took that idea further than I expected. It reads every message across , most of which have no human name. Where a persona holds a dozen facts, the model picks up tone, pacing, hesitation, and patterns we never thought to define. It works like a therapist who reads body language and tone of voice, not just the words, and catalogs all of it.

That changed my job. Instead of defining the user in a document, I could choose what the model should watch for: self-awareness, agency, resilience. We start with human-defined we can understand and check, and the model reads them in every conversation, at a depth no research team could match.

Impact: Everything below builds from this. The classifiers read the signals. The scaffolding uses them to keep a new user at a level they can handle instead of opening with hard questions. And the feedback loop tunes what the model watches for as we learn.

Slide: LLMs Deepen Who We Design For — three columns. Persona, the profile of a user: age, role, scenario, goal. Archetype, the user’s state in the product: power user, creates own shortcuts, pride in learned skills, a beginner three months ago, ready for harder questions. Dimensions, the signals the model picks up: shorter replies at night, momentum after naming an issue, slight signs of fatigue this week, hesitation rising now. Adapted from “Designing Complex Apps for Specialized Domains”, Kate Kaplan, Nielsen Norman Group (2020).

Nielsen Norman’s archetypes: a single read of user state. The product’s does the same thing at scale — reading the in every message across .

Meeting Users Where They Are


Three examples of how the product opens with different kinds of user. Two elements do the work: the Loading Screen Messages (shown on the loading screen, they set the tone before a word is exchanged) and the Opening Chat Prompts (HealthBot’s first move into conversation).

Kyler Z.

Basic Onboarding — Kyler

Short, low-pressure framing. No vocabulary the user has to look up.

Loading Screen Messages

  • Can’t sleep? Give HealthBot a try.
  • Need to ‘brain dump’? HealthBot is ready to listen.
  • Stuck in your thoughts? We can unpack them.
  • Relax, you are now in a judgment-free safe place.
  • There’s no right way to use HealthBot, but if you improve or feel better, it worked.

Opening Chat Prompts

  • No pressure to have a point. What’s going on today?
  • We can keep it easy. How’s the day treating you?
  • Rough shift, or just killing time? Either’s fine.

Light, low-commitment, no assumptions — he’s new.

Jakob B.

Deep Engagement — Jakob

Belief-and-mindset framing for the user who is already past the surface.

Loading Screen Messages

  • Small things done over and over add up. Keep showing up.
  • To change anything, you must first change your mind.
  • What you think is what you become.
  • The solution doesn’t have to be perfect; it just has to help.

Opening Chat Prompts

  • You’ve been showing up. What’s been different lately?
  • What’s something you’re trying to work through right now?
  • What’s one thing you’d want to be different a month from now?

Reflective, invites depth without being corny — he’s engaged but guarded.

Evelyn R.

Retention — Evelyn

Identity-affirming language. The job here is to remind the user why they keep coming back.

Loading Screen Messages

  • You are the expert of your own life.
  • It starts with owning your own story.
  • You are here because you know you are important.
  • What was once brick by brick is now a view of your skyline.

Opening Chat Prompts

  • Good to have you back. This one’s still just for you. How are you, really?
  • You spend all day holding things up for other people. What’s here for you today?
  • You keep coming back for a reason. What’s pulling at you today?

Warm, knowing, gently turns the lens back to her — she’s returning and puts herself last.

Three intros, same product, three different reads on what the user actually needs in the first 30 seconds.

Incremental Steps with Scaffolding


The model’s starts every topic light and only goes deeper when the user shows they can handle more. If they don’t, it stays light.

  • Self-awareness (Pilot)Noticing, naming, reflecting, meaning-making
  • AgencySmall choices, boundaries, decisions, life direction
  • Emotional regulationRecognizing a feeling, sitting with it, responding instead of reacting
  • ConnectionAcknowledging isolation, reaching out, sustaining relationships
  • ResilienceSetback, recovery, reframing, growth after hardship

Impact: Self-awareness is the dimension I built out by hand (see right), so the same structure can be duplicated for any other dimension worth improving.

Not a Chatbot, an Agent


The words get used interchangeably. The difference decides everything about how this product behaves, because its users routinely say “I’m fine” when they’re not.

  • ChatbotReacts to what someone said. The message comes in, a response goes out. The surface of the words is all it has.
  • AgentReads, decides, and acts on what it perceives. It holds what it knows about the person, weighs what’s beneath the message, chooses a move, and sometimes the move is to hold back.

A respectful conversation is one that can aid self-discovery and self-reflection — which means the agent has to know what someone’s been working through, not start from zero every message.

That’s what I pushed for: a history people can look back on, goals they can return to. The original product had none of it. Every reply came from the last message alone, and leadership believed remembering a user’s circumstances would take away their agency. Without memory, this was an open journal disguised as a chatbot.

Impact: “Memory” became its own three-person team.

Two paths from the same message. On the left the chatbot runs Interface, Intent Matching, Business Logic and a scripted response; on the right the agent runs Interface, Perception, Memory, Decision, Actions and Guardrails

The chatbot matches a message. The agent reads a person.
See the Conversational AI Test Bench

Service Blueprints


I routinely met with each ML engineer one on one. Everyone was working alone: services duplicated each other, handoffs broke daily, nobody owned the failures.

And the interviews surfaced the real finding — everyone was independently building toward the same unnamed thing, an agent. One engineer’s “conductor” was another’s “switch” or “mothership.”

Impact: The blueprints did two jobs. They showed us what we actually had: every service, what it did, who owned it.

And they gave us the 90-day cleanup plan, setting the stage for in-chat classifiers.

Service blueprint, current workflow
Current workflow The current state as I found it: ten steps from user message to reply, seven services, and no one who could draw it. I built this map from one-on-ones with each engineer, naming every service, what it does, and what it hands to what. The named services existed; the picture of them working as one system did not.

Mapping it changed the conversations. Redundancies became visible, ownership gaps had nowhere to hide, and the team saw for the first time that they were all building parts of the same thing: an agent. What the map also made plain: nothing listened. Every turn started from zero, and everything the user revealed was gone by the next message.

Classifiers + Understanding Pipeline


The classifiers are a team, not a feature. Nine specialized ‘bots’, each with one job: reading atomic traits, tone, risk and safety, resilience, insights, narrative arcs, goals and motivations.

Each bot has a goal, a scope, and example insights it’s expected to produce, and they run in phases of increasing complexity, from granular signal detection up through narrative modeling.

The understanding pipeline: detection, indication, identification, classification and re-evaluation across a single conversational turn
The Understanding Pipeline What the bots do with a turn. A user says “Yep, I’m OK. So what now?” and the system runs it through five moves: detect the surface signals, form a hypothesis, build confidence in a trait, classify it into an organized model of the user, and re-evaluate as life changes. Understanding the user became its own pipeline, with its own architecture, not a side effect of generating a reply.

How We Guide the Model,
How the Model Guides the User


The product is the model’s behavior. Model design and product strategy both guide it.

  • Model GoalHelp users self-reflect through conversational support.

In and Out of Model’s Scope

Category Teach Avoid
Therapeutic frameworks basics Clinical diagnosis
Emotions Basic emotions, self-awareness Telling users what they feel
Relationships Self, work, interests Prescribing what to do
Finance Thinking about spending Saving advice, financial planning
Safety Crisis awareness, de-escalation Handling the crisis itself
Religion Only if user chooses Guiding beliefs in any direction

The pattern across every row: teach the skill, never perform it for them. The model builds capability, not dependency.

Impact: One strategy at four levels. The company names the individual’s journey. The product names the destination: resilient users. The model names the vehicle: self-reflection. The scaffolding names the stages, starting at self-awareness because nothing else works without it.

Ideation and Study Decks

HealthBot UX Concepts deck cover
Concept directions catalogued as they came up, before anything earned a place in the backlog.
User Maturity Concepts deck cover
A study of how far users can read and regulate their own state, and the levels in between.
HealthBot User Testing deck cover
Rounds of user testing across the build, from early concepts through the flows that shipped.

Conversational Signal Processing


An ML notebook that turns open conversation into structured data the system can act on: how to pace the conversation, when to offer support, when to acknowledge growth. It reads a user’s language for behavioral signals — emotional cues, cognitive distortions, resilience markers, motivational patterns — each detected by its own narrowly scoped classifier.

The classifiers increase in complexity:

  • Granular Signal DetectionAtomic traits, tone
  • Micro-Pattern DetectionRisk and safety, self-recognition, resilience
  • Emergent InsightsInsights and reflections
  • Narrative ModelingStory arcs, breakthroughs, goals, motivations

Impact: Each reports what it observes in the user’s language, with a score and a sample insight; the composite is what the decision layer acts on.

AI Response Generation


A notebook that takes Phase I’s output and shows how a reply gets built from what is underneath a message — emotional intensity, coping style, where the person is in their story — instead of its surface words.

Each response combines four parts:

  • Empathic ReflectionReflect the emotional strain
  • ValidationNormalize the experience
  • Cognitive ReframingA gentle alternate perspective, without invalidating the user’s truth
  • Forward GuidancePoint toward a small next step

Impact: A CBT-aligned reply built from behavioral evidence rather than keywords. Not diagnosis or advice, but the kind of thing a thoughtful friend with therapeutic training might say.

Reinforcement Learning from Human Feedback (RLHF)


The grading interface collects human feedback with a thumbs up or down beside each response. It’s the same approach major AI labs use to teach models what good responses look like.

Before it, model comparison was manual: the same prompts run against different models and scored by hand. Comparative model tests (XLSX)

Therapist review console: past sessions listed on the left, the conversation in the middle with an expandable reasoning trace, and a review panel on the right scoring HealthBot’s emotional response, with notes, a safety checkbox and flag for group review View the Therapist Review Console

Therapists work through past sessions, scoring each response for overall quality and SFBT adherence, adding notes, and flagging the ones that need a second look.

Impact: Therapists tune the model’s behavior, one graded response at a time, and at scale those judgments become how it responds.

What Having Memory, Classifiers, and Understanding Actually Changes


Examples of a different model behavior over a . Italics above each turn is what the model is recognizing or doing.

Evelyn Evelyn — opening message “I had a rough night. Mom didn't sleep so I didn't sleep. Now I'm at work and I can't focus.”
Original
Turn 1 · Reactive question
How does that make you feel?
Tired. Frustrated. Like I'm not doing anything well right now.
Turn 2 · Loop deeper
What does it make you think about, to feel tired and frustrated?
I don't know. That I'm failing, I guess. At all of it.
Turn 3 · Loop again
When you think about feeling like you're failing — what does that thought make you feel?
Worse. This isn't helping.
Turn 4 · Loop until user drops
What is it about thinking about feeling worse that brings up for you?
User has stopped responding.
Adding Memory + Classifiers
Turn 1 · Acknowledge before asking
That's a hard combination — no sleep and a workday with no margin. The fact that you're here, on a day like this, is worth noting.
Tired. Frustrated. Like I'm not doing anything well right now.
Turn 2 · Use memory + surface counter-evidence
You mentioned a couple of weeks ago that the nights with your mom were starting to compound. Sounds like that's still where the weight is sitting. When you say ‘not doing anything well’ — is that the whole day, or are there a couple of things you'd actually rate as fine?
I mean. I got both kids out the door on time. I covered for someone at work who couldn't come in.
Turn 3 · Reframe with user's own evidence
Both of those count. You're describing a day where you showed up for three different people before you got to think about yourself. That's not ‘not doing anything well.’ That's depleted, which is a different problem with a different solution.
Yeah. I guess that's true.
Turn 4 · Link past insight + propose small action
Last time we talked, you said the hardest part was the guilt at the end of the day when nothing got done for you. If today ends without something small for you in it, the guilt's going to compound the depletion. What's a 10-minute thing today that would belong to you?

What memory and classifiers did per turn — including what it chose not to say

PerceivedDecidedSuppressed
Turn 1 Caregiver fatigue, no margin Acknowledge before extracting Clinical-sounding questions (“rate your stress 1–10”)
Turn 2 Self-criticism loop activating (“not doing anything well”) Bring in prior session memory + ask for counter-evidence “Have you considered respite care for your mother” — too solution-forward, user is in processing mode
Turn 3 User supplied counter-evidence she didn't realize she had Reflect it back and reframe depletion vs. failure Praise that would feel hollow
Turn 4 User accepted the reframe Connect to a prior insight (guilt at day's end) and propose one small concrete action A list of options (would feel like homework)

What the Model Gets Right, What It Costs


Most of what an evaluation needs is already built: a test set, a , human grades. What’s missing is running everything as a batch — every scenario scored, cost counted per conversation. That’s the , a separate tool now in progress.

Industry Term Project Artifact Status
45 user test messages (synthetic) Built: Test Bench (opens in a new tab)
“In and Out of Model’s Scope” Table Built: Model Design
Therapists grading responses (up, down, notes) Built: Human-in-the-Loop
Governance Toggle (same message run, ruleset on and off) Built: Test Bench (opens in a new tab)
Emergency Scenarios (must route to human help) In Progress
Feedback panel collecting grades from live runs In Progress
Full scenario battery run as a batch, scored Planned
Same battery across model tiers, quality against cost Planned

Try HealthBot


The test bench is a dual-pane view built for demoing the inputs for the entire agentic pipeline, not just chatting with it.

Demo the Test Bench

Live and interactive, running as a Node server against the Anthropic API — Demo the HealthBot Test Bench

The Latent Expert


A future feature (unshipped): the user can opt into a framework they want to grow within — a mythology, a philosophy, a faith. The agent learns it and reflects through it, but never instructs. The user works things out inside a way of thinking they chose; the agent just knows it well enough to keep them company there.

It’s opt-in and set in settings, never inferred from conversation. Something this personal has to be chosen, not switched on quietly.

Kyler

Kyler

Following the Hero’s Journey

Drawn to Star Wars or Marvel, Kyler opts into mythic-narrative reflection.

The agent never casts him as the hero or narrates his arc. It quietly frames things the way that structure does — the ordeal, the return — so a hard week reads as a chapter, not a conclusion.

Impact: Kyler recognizes the pattern himself; the agent just holds the shape.

The agency stays his.

Evelyn

Evelyn

Exploring a Wisdom Tradition

Curious about a tradition, or rooted in one, Evelyn opts into it.

The agent becomes fluent in that framework’s language of meaning, duty, and grace, and lets it inform how it reflects — offering the tradition’s questions, not its answers.

Impact: Whether Evelyn is new to the tradition or rooted in it, she is met in her own idiom and never preached to.

The agency stays hers.

Jakob

Jakob

Focusing on a Philosophy or Practice

Stoicism, a coaching discipline, a way of thinking Jakob admires — he opts into it.

The agent reflects through that discipline’s habits (what’s in your control, what story you’re telling) without lecturing the doctrine.

Impact: The framework becomes the grain of Jakob’s conversation, not its content.

The agency stays his.