I just found out that OpenAI publicly announced ChatGPT Health on January 7, 2026, and that, initially, availability will be gradual through a waitlist.

What matters isn't just "a new feature," but the concept: an experience within ChatGPT designed specifically for health and wellness, with more context when the user chooses to connect personal information (medical records and health apps). OpenAI presents it as a tool to be better informed and prepared for clinical conversations, without replacing health professionals.

What exactly is ChatGPT Health?

ChatGPT Health is a "dedicated" experience within ChatGPT for health conversations. Its differentiator is that it offers a separate space ("Health") with:

  • Chats isolated from the rest
  • Health-specific memory
  • Optional connection of data sources, such as medical records and health apps

One key point: OpenAI states that health chats, files and memories are not used to train its base models. That's not a minor detail.

HealthBench: the piece that explains "how it improves" without using real user data

To understand the approach, we need to talk about HealthBench and how it fits into this story.

HealthBench is an open health benchmark with 5,000 multi-turn conversations (chat-style) that simulate realistic scenarios. Each conversation is scored using rubrics written by physicians (262 physicians took part in creating the criteria) and tens of thousands of unique criteria that score aspects such as:

  • Safety
  • Clarity
  • Handling uncertainty
  • Communication
  • When to escalate to professional care

Unlike "exam questions," HealthBench aims to measure response quality in context, the way a clinician would judge it in real life.

In practical terms: OpenAI uses HealthBench as an evaluation (not as a source of user data) to sustain a "quality system" that lets the model improve without relying on real clinical data from users.

Current scope: what is it actually good for?

According to what has been shared, ChatGPT Health is being rolled out for cases such as:

  • Summarizing clinical information provided by the user (e.g., PDFs, lab results, visit summaries).
  • Explaining results and trends "in plain language" and helping prepare questions for a medical appointment.
  • Identifying longitudinal patterns (habits, sleep, activity, nutrition) when the user connects apps/wearables.

Limitations: what shouldn't it be used for?

Important limits are also highlighted:

  • Not for diagnosis or treatment.
  • For now, only available in: the US, EEA/EEE, Switzerland and the UK.
  • EHR integration: US only, and for users over 18.
  • Like any LLM, it can get things wrong due to missing data or misinterpretation; that's why it should be seen as support, not "clinical authority."
  • As a consumer tool, it isn't the same as a regulated clinical system; caution is warranted with sensitive data.

Potential benefits: what could actually change

Personally, I think the most transformative part isn't "summarizing PDFs," but its cultural impact:

  • It can empower patients to understand their own information.
  • It could reduce friction and improve the quality of conversations with health professionals.
  • It could add continuity by compiling a patient's scattered data, even when spread across multiple systems.

Interoperability: this is where FHIR comes in strong

I couldn't leave out interoperability. And there's good news: in the version enabled in the US, there's mention of integration with FHIR-based data repositories through partnerships with providers to enable access under:

  • BAA (Business Associate Agreement)
  • Identity
  • Consent
  • Auditing

All under consumer-mediated exchange models: the patient authorizes and mediates access.

If this holds up, it opens up a really interesting conversation: how a conversational assistant integrates into the clinical ecosystem without "skipping" governance, consent and traceability.

Real risk: could this increase self-diagnosis or self-medication?

We need to pay attention to possible counterproductive effects. Yes, there's a risk that self-diagnosis and self-treatment could increase among certain profiles, but it's not an automatic or inevitable effect.

It depends heavily on:

  • How it's used
  • How robust the "guardrails" are
  • The context of the health system (appointment access, costs, wait times)
  • The local culture around self-medication

Next step

I've already asked to be added to OpenAI's waitlist, since I'm interested in testing it with FHIR repositories integrated with EHR systems and documenting a PoC-based experience.

If it's enabled, I'll share practical findings: what works, what doesn't, and what implications it has for interoperability, governance and patient experience.