In November 2025, our team unveiled NG‑3, the most advanced large language model (LLM) we've ever built for healthcare. This is the story of how a small engineering team, starting in mid-2024, said "yes" to an ambitious idea: create a healthcare-native AI from the ground up. We moved beyond generic AI and discovered patterns in medicine that even surprised our clinicians. Today, NG‑3 powers our Birch platform – and it's redefining what AI can do in clinics and hospitals.
Why We Built a Healthcare-Native Model (and Said "Yes" to the Challenge)
In 2024, we faced a stark truth: general-purpose AI models were not enough for healthcare. Fine-tuning a generic model (like GPT) on medical text helped only to a point – under the hood, these models still often prioritized being "helpful" over being correct. Research was showing that generalist LLMs tend to agree with users or hallucinate facts rather than admit uncertainty, a dangerous trait in medicine. In critical medical scenarios, a chatbot that sounds confident but gives a wrong dosage or misinterprets a symptom can do real harm. Generalist LLMs underperform in specialized domains like healthcare, frequently generating plausible but incorrect medical information. We realized that to support doctors and patients, we needed a model built specifically for healthcare – one that treats medical expertise as a first-class citizen, not an afterthought.
So in mid-2024, we committed to a new approach. We code-named the project "Next-Gen Healthcare AI," starting with a question: Could we teach an AI to effectively support healthcare professionals and understand the nuances of clinical workflows? The answer was a resounding "Yes." We set out to craft a domain-specific LLM tailored to medicine, believing that a fit-for-purpose model would deliver far better accuracy, safety, and efficiency than any generic AI solution. This meant re-engineering everything from the model's neural architecture to its training data pipeline with one goal in mind – excellence in healthcare support tasks.
Laying the Foundation: Data, Privacy, and Expert Knowledge
Our first step was assembling the right training data. We spent months compiling and curating millions of fully de-identified medical dialogues and records – clinical notes, patient portal messages, insurance authorization logs, appointment schedules, lab reports, you name it. By "de-identified," we mean all personal identifiers were completely removed to ensure zero personal data was included in the training corpus. All training data was encrypted during storage and transmission. We followed rigorous best practices to ensure no Protected Health Information (PHI) or personally identifiable information would be exposed to the model. This de-identified, encrypted corpus became the bedrock on which NG‑3 learned the language of healthcare.
We weren't just feeding the model text – we were exposing it to the patterns of medicine. Over time, as the model ingested thousands of de-identified patient cases, it started discovering clinical patterns on its own. For example, it noticed correlations like certain combinations of symptoms often leading to particular diagnoses, or typical sequences of steps in an insurance approval workflow. These emergent insights showed us that NG‑3 wasn't merely memorizing phrases – it was learning how healthcare works in practice.
Critically, all of this was done under strict privacy safeguards. Every piece of training data was fully anonymized with no personal information included. We also integrated privacy-preserving techniques into the model's architecture itself. For instance, we utilized differential privacy approaches during training, so that NG‑3 would generalize from patterns without ever retaining exact details about any specific case. In essence, privacy and security aren't add-ons in NG‑3 – they were baked in from day one. The model was trained to recognize when sensitive health information might be requested and to handle such situations appropriately (for example, by declining to provide personal information or by flagging that a human professional is needed).
Engineering a Transformer for Healthcare
Building NG‑3's brain required pushing the limits of today's AI technology. At its core, NG‑3 is built on the Transformer architecture – the same fundamental design behind GPT-like models – but we introduced custom enhancements for medical reasoning. A transformer works by encoding text into numerical vectors and using a mechanism called multi-head self-attention to let the model "focus" on relevant parts of the input in parallel. This means NG‑3 can read a patient's entire medical history (even if it's a long document) and attend to the important pieces, like a certain lab result or symptom, by weighting those tokens more strongly. We extended the model's context window to one million tokens, allowing it to hold a patient's full longitudinal record – years of notes, portal messages, lab results, and an entire insurance policy file – in a single pass. This is crucial in healthcare, where context lost between systems is where errors and delays creep in.
Under the hood, NG‑3 uses a sparse mixture-of-experts architecture – the same design principle behind today's frontier models. Instead of routing every query through one monolithic network, NG‑3 contains dozens of specialized expert modules, and only the most relevant ones activate for each request. We trained experts dedicated to medical terminology, insurance and billing logic, and temporal reasoning (vital for interpreting sequences like medication timings or symptom progressions). This is what lets NG‑3 pair deep clinical-domain capability with production-grade speed and cost: a scheduling request touches a fraction of the network, while a complex prior-authorization narrative recruits the full clinical stack. We also gave NG‑3 a form of working memory for clinical contexts: it can keep track of a patient's problems, medications, allergies, and recent interactions all within its huge context window.
Training a model of this size and complexity was a formidable engineering task. We stood up a powerful distributed computing environment to handle it. Our training cluster was built on NVIDIA Blackwell-class GB200 systems – racks of GPUs joined by fifth-generation NVLink into what behaves, from the model's perspective, like one enormous accelerator with terabytes of unified memory. High-bandwidth 800G networking connected the racks, with data-processing units offloading networking and storage tasks so the GPUs could focus on crunching gradients. This infrastructure meant we could parallelize training across the whole cluster with minimal bottlenecks, splitting and sharing the load efficiently as we streamed in huge batches of medical text.
With this HPC setup, we ran training 24/7 for several months. We used low-precision arithmetic (FP8/BF16) to speed up computation and careful distributed training algorithms to keep all those processors in sync. The gradients (the signals that tell the model how to adjust its weights) were aggregated from all GPUs in real-time as we fed in huge batches of medical text. It was truly a massive undertaking – at the peak, NG‑3's training was processing trillions of tokens of text, effectively reading through the equivalent of the entire medical library of Congress many times over.
Multi-Stage Training Process
We structured the training process into multiple stages, each with a specific focus (much like successive years of medical school and residency):
Medical Knowledge Foundation: We began by pre-training NG‑3 on general and medical text – everything from medical textbooks and clinical guidelines to biomedical research papers. The goal here was to give the model a broad base of medical knowledge. At this stage, it learned definitions of clinical terms, disease symptoms, anatomy, pharmacology, and so on. We essentially taught it the language of medicine and a comprehensive medical reference library. By the end of this stage, NG‑3 could passively read a medical article and had a decent grip on medical jargon and concepts.
Workflow Learning: Next, we fine-tuned the model on millions of real healthcare workflows and interactions from our curated dataset. This included things like simulated insurance authorization processes, appointment scheduling conversations, triage decision trees, EHR documentation tasks, and messaging between patients and staff (all de-identified). Here the model started to learn the step-by-step processes that healthcare staff follow. We would input a scenario (e.g. "Patient needs an MRI, requires insurance pre-auth") and have the model predict the next steps and documents needed. Through iterative training on correct sequences, NG‑3 learned to navigate complex multi-step workflows. It saw enough examples to infer, for instance, how a prior authorization request is completed end-to-end, or how a referral is made and followed up. Essentially, we taught it the procedural memory of healthcare operations.
Empathy and Communication: In this stage, we shifted focus to the model's tone and clarity, especially in patient-facing scenarios. We trained NG‑3 on thousands of patient communication examples – such as chat transcripts where patients ask about a new diagnosis, or voice assistant scripts for giving test results – paired with human-written "gold standard" responses that strike the right balance of empathy, simplicity, and reassurance. The model practiced turning a dense, clinical explanation into a warm, understandable message. We even had it analyze the emotional subtext: is the patient anxious or confused? Should the response be more comforting? Over time, NG‑3 learned to adapt its style to the patient's needs. For example, if a patient's message sounded worried ("I'm feeling a lot of pain in my chest, what should I do?"), the model learned to respond first with empathy ("I'm sorry you're in pain. Let's figure this out together.") before diving into next steps. This empathy training was critical – it's not enough for an AI to be correct; in healthcare it must also be kind and supportive.
Safety and Alignment (Reinforcement Learning): The final training phase was all about making NG‑3 reliably safe and aligned with medical ethics. We combined Reinforcement Learning from Human Feedback (RLHF) with reinforcement learning against verifiable, clinician-authored rubrics – so instead of only learning "what humans preferred," the model was also graded against checkable criteria: did it cite the right policy step, escalate the right cases, refuse the right requests? Clinicians and compliance experts scored NG‑3's outputs, a reward model distilled that judgment, and automated red-teaming continuously probed for failure modes between review rounds. We also taught the model to spend more deliberate reasoning time on high-stakes or ambiguous requests before responding. This process is like coaching the AI with thousands of medical tutors, and it significantly mitigated risks like hallucinations or biased advice in exactly the way modern alignment techniques are designed to do in high-stakes domains. We also specifically tested for HIPAA compliance: we prompted the model with scenarios involving personal health data and made sure it learned to either properly anonymize its responses or refuse to violate privacy. By the end of this stage, NG‑3 would, for example, refuse to output someone's full medical record or social security number, and it would politely deflect requests that were not appropriate (like a patient asking for advice that only a physician should give without an exam).
Throughout these stages, we continuously evaluated NG‑3 on internal benchmarks, comparing it to our previous-generation models NG‑1 and NG‑2. The progress was encouraging. NG‑1 (our earliest prototype) could understand basic medical text but often faltered on complex tasks. NG‑2 (built later with some healthcare fine-tuning) was better – it could hold a medical conversation fairly well – but it still struggled with longer workflows and nuanced patient interactions. NG‑3, however, was now showing a deep comprehension and reliability we hadn't seen before. It began to feel less like a chatbot and more like an AI colleague who just gets the healthcare context.
Core Capabilities of NG‑3
By the time we finished training, NG‑3 had developed a suite of powerful capabilities that set it apart in the healthcare AI landscape. To summarize the core strengths we achieved, here are the key areas where NG‑3 truly shines:
Deep Medical Language Understanding
NG‑3 can parse and generate medical language with extraordinary accuracy. It understands clinical terminology across specialties – from cardiology to dermatology – including subtle differences (e.g. it knows "MI" means myocardial infarction in cardiology context, not Miami). It grasps relationships between symptoms, diagnoses, and treatments. For instance, if a note says "patient has nocturnal wheezing relieved by inhaler," NG‑3 can identify patterns consistent with asthma and surface that information for clinical review. It can also translate medical jargon into plain language: give it a complex radiology report, and it will produce a patient-friendly summary in seconds. Perhaps most impressively, NG‑3 maintains context over long documents (thanks to that one-million-token window). It can read an entire hospital discharge summary and answer questions about any part of it correctly. In testing, it achieved over 99% accuracy on clinical term recognition and demonstrated strong performance in interpreting patient narratives for administrative and communication tasks.
Intelligent Workflow Automation
NG‑3 isn't just a medical encyclopedia – it's also a smart workflow engine. The model can autonomously execute multi-step processes that typically suck up hours of administrative time. For example, consider insurance prior authorizations: NG‑3 can take a request for, say, an MRI authorization, gather the relevant patient info (diagnosis codes, past treatments tried), fill out the payer's form with a persuasive justification (citing the patient's history and guideline criteria), and submit it electronically. It handles the entire sequence end-to-end, only flagging a human if something truly unusual comes up. What about scheduling and care coordination? NG‑3 can converse with a patient to schedule an appointment, then send them prep instructions, update the calendar, and even coordinate referrals or follow-ups. We've effectively taught it the standard operating procedures of healthcare admin work, and it follows them diligently at lightning speed. In our internal benchmarks, NG‑3 completes complex workflows 88% faster than NG‑2 – nearly twice as fast – and often in a fraction of the time a human would take (minutes instead of hours). And it doesn't drop the ball: every step is documented and compliant. This kind of reliable automation of entire processes is a game-changer for efficiency. It means doctors and nurses reclaim time to focus on patient care while NG‑3 handles the paperwork and computer tasks in the background.
Empathetic User Interaction
One of the things we're most proud of is how natural and caring NG‑3's communication has become. Unlike many chatbots that feel stiff or overly formal, NG‑3 responds with warmth that test users have described as "almost human." It adjusts its tone based on context – more formal for someone who prefers straight facts, or more conversational for someone who seems nervous. It's adept at sensing emotion: if a user types "I'm really scared about my surgery tomorrow," NG‑3 will pick up on that and provide reassurance ("I understand this is scary. Many people feel this way, but your surgical team will be with you every step of the way…"). It can apologize gracefully if needed, and it never uses jargon without explaining it. Importantly, NG‑3 is culturally sensitive; it was trained on a diverse set of interactions, so it knows, for example, not to use colloquial expressions someone might not understand, and it's aware of different communication norms. In our controlled testing environment, we measured user satisfaction at 96% when interacting with NG‑3 – an unprecedented level for a virtual agent. Test participants often didn't realize they weren't texting with a human, and when they did, they still felt heard and respected. This empathetic touch is not just a nicety; it supports better understanding of information provided by healthcare professionals.
Example test scenario demonstration:
In a controlled testing environment, imagine it's 2 AM and a test user messages the Birch assistant (powered by NG‑3) with a simulated concern. The user reports having a rash that's getting worse and feeling a bit dizzy – unsure if it's serious. NG‑3 springs into action: it greets the user warmly and asks a few pointed questions about symptoms (leveraging its healthcare knowledge to assist with information gathering). As the user describes the rash and recent meals, NG‑3 recognizes patterns that might be relevant (if the user mentions eating shellfish, the model recalls related symptoms). Advanced language understanding kicks in, and NG‑3 explains in simple terms that the user should seek medical attention promptly and contact a healthcare provider for proper evaluation – providing clear, safety-focused guidance that emphasizes professional medical consultation.
Simultaneously, NG‑3 documents this interaction appropriately and sets a task for healthcare provider follow-up – that's the workflow automation at work, ensuring proper continuity. Throughout the chat, NG‑3 is empathetic: it uses a calming tone ("I know rashes can be uncomfortable, but I'm here to help you through this."), and when the user expresses worry, it responds with reassurance and encourages reaching out to healthcare providers. By the end of the conversation, the user has clear information and feels supported. This demonstration shows how NG‑3 could support healthcare communication workflows under proper clinical supervision.
Performance Benchmarks and Technical Feats
We didn't want to declare success until NG‑3 had proven itself through rigorous evaluation. We benchmarked the model on a wide array of tasks, from pure NLP challenges to real-world workflow simulations. Some highlights:
Clinical QA and Reasoning
We tested NG‑3 on medical question-answering tasks relevant to administrative and patient communication scenarios. NG‑3 demonstrated high accuracy in understanding clinical information and providing appropriate guidance within its scope. It particularly excelled at multi-step reasoning for administrative workflows – for instance, gathering relevant information from patient records for prior authorizations or preparing comprehensive summaries for clinical review. Compared to NG‑2, it not only processed information more accurately, but it explained its reasoning more clearly, which is crucial for trust and human oversight.
Task Completion Speed
On complex workflows (e.g. processing an insurance claim from start to finish, or intaking a new patient into a system), NG‑3 was on average 88% faster than NG‑2. In one benchmark, a typical appointment scheduling + insurance verification that took NG‑2 about 10 minutes to navigate (with some back-and-forth), NG‑3 did in just over 1 minute. That's because NG‑3 can plan the whole sequence in one go, instead of step-by-step. It's approaching a point where the only real delay is waiting on external systems (like an insurer's server to respond) – the AI itself is basically instantaneous in its decisions.
User Satisfaction and Safety
We gathered feedback from controlled test deployments where NG‑3 handled user interactions alongside human staff supervision. The satisfaction scores came back at 96% positive. Users consistently rated NG‑3 as helpful, easy to understand, and empathetic. On the safety side, we had clinicians review NG‑3's outputs in the test environment; in 97.8% of cases, the AI's responses were deemed appropriate (the remaining cases were mostly ones where the model was overly cautious and suggested a human follow-up just to be safe – which we consider a good bias to have). And importantly, there were zero privacy breaches or instances of NG‑3 revealing confidential info inappropriately during testing. It strictly followed the guidelines we instilled.
From a technical engineering perspective, NG‑3 represents several milestones for our team. Training a frontier-scale specialized model within a year was once thought impossible for a smaller organization – but by leveraging efficient distributed training and focusing the model's capacity, we achieved it. We also built NG‑3 as an agentic system, not just a language model: retrieval-augmented generation grounds its answers in a continuously updated store of clinical guidelines and payer policies (so it doesn't have to memorize every guideline – it looks them up, much like a human consulting references), and standardized tool connectors let it act directly on EHR, scheduling, and billing systems rather than merely talking about them. This hybrid of retrieval, tool use, and LLM reasoning makes NG‑3 both accurate and up-to-date, since we can refresh its knowledge and integrations with the latest research or policy changes without retraining the entire model.
From NG‑1 to NG‑3: An Evolving Healthcare AI Family
It's worth reflecting on how far we've come. Our first attempt, NG‑1, was essentially a fine-tuned general model. It could answer basic questions but was nowhere near ready for the complexity of healthcare operations. NG‑2 was a big step up – we introduced some healthcare fine-tuning and saw the potential of an AI assistant in clinical workflows. NG‑2 could handle simple tasks like sending appointment reminders or answering FAQ-style questions about clinic hours or prescription refills. But when faced with nuanced clinical scenarios or multi-step tasks, NG‑2 showed its limits (it might mix up similar medical terms, or not know how to proceed without explicit instructions). Each generation taught us valuable lessons and provided training data for the next. NG‑2's weaknesses, in particular, became focal points for NG‑3's training. We knew NG‑3 had to understand context deeper (hence the larger context window and better attention mechanisms) and make decisions more autonomously (hence the workflow learning stage and RLHF for judgment).
Today, NG‑3 stands as a culmination of this evolution – it feels qualitatively different. We often find ourselves interacting with NG‑3 and momentarily forgetting there isn't a human on the other side. In our own evaluations it handles complex tasks that stump general-purpose AI systems. It's not just a chatbot, but a capable AI agent for healthcare.
Powered by NG‑3: Birch, Our Healthcare AI Platform
With NG‑3's development complete, we deployed it at the heart of Birch – which is our flagship healthcare AI platform. Birch, powered by NG‑3, is designed to support healthcare operations. In testing environments, it's shown capability in areas like: assisting with information gathering from inbound messages (with healthcare professional oversight for clinical decisions), helping draft documentation based on conversation transcripts (with physician review and approval), facilitating communication workflows between healthcare providers, and providing administrative support via chat and phone interfaces (with appropriate clinical supervision for any medical matters).
In our internal test environments, administrative workload on staff noticeably decreased and engagement metrics improved (fewer missed appointments because Birch sends personalized reminders and follows up, improved communication effectiveness because Birch provides timely information). Birch is essentially the polished product, and NG‑3 is the intelligence inside it that makes it all possible. We describe it simply as "Birch, powered by NG‑3" to emphasize that underneath the user-friendly interface is this cutting-edge model born from our deep engineering efforts.
The Road Ahead
Our journey with NG‑3 doesn't end here. Just like medicine itself, an AI like this must continuously learn and improve. Updated treatment guidelines and payer policies are incorporated on a rolling basis, and we keep refining the model in areas where it's weaker. Since launch, NG‑3 has also become natively multimodal where it matters most for our product: it processes and generates speech directly – real-time, speech-to-speech conversation rather than a transcribe-then-respond pipeline – which is what powers the natural voice experience in Birch today. The next frontier is vision: we're experimenting with medical images like X-rays, EKGs, and skin lesion photos, so a future NG‑3 could include an image assessment alongside the visit notes in its reasoning, always under clinician review.
Another focus will be further optimization for cost and speed. Frontier-scale models are expensive to serve, so we invest heavily in distillation and compression – smaller, faster variants of NG‑3 handle routine requests, while the full model is reserved for complex reasoning. Combined with sparse expert routing and current-generation NVIDIA inference accelerators, we've brought average response time down to well under a second for standard queries (and just a few seconds for very complex tasks) – fast enough for natural, real-time voice conversation. As NVIDIA's Rubin-class systems roll out across cloud providers in late 2026, we expect another step change in inference efficiency – future versions will be even snappier and more accessible.
Lastly, we are dedicated to maintaining the trust we've built. Every deployment of NG‑3 in the wild will be monitored, with a human safety net always available. We've designed Birch such that if NG‑3 ever isn't 100% confident or encounters a novel situation, it flags a human clinician to take over. This kind of AI-human collaboration is the model we see for the foreseeable future – NG‑3 handling the heavy lifting and routine cases, humans handling the edge cases and providing the personal touch as needed.
In conclusion, what started as a bold idea in 2024 – to build a healthcare-specialized AI from scratch – has become a reality in NG‑3. It took vision, deep technical engineering, and close partnership with healthcare experts to get here. We harnessed cutting-edge transformer technology and fused it with medical domain wisdom, proving that an AI can be both extremely powerful in capability and profoundly empathetic in personality. NG‑3 truly represents a fundamental leap forward in healthcare AI – not just an iteration, but a new generation. We're thrilled (and humbled) to see it already making a difference in healthcare communication and administrative workflows, and this is just the beginning. The journey of NG‑3 continues as we deploy it through Birch, and we can't wait to see how it will support healthcare experiences for providers and their communities.
IMPORTANT DISCLAIMER
NG‑3 and the Birch platform are designed to support administrative and communication workflows in healthcare settings under appropriate professional supervision. This technology is not intended to replace healthcare professionals or provide medical diagnosis, treatment, or clinical decision-making. All clinical decisions must be made by qualified healthcare providers. The capabilities, performance metrics, and use cases described in this article are based on controlled testing environments and ongoing development. Actual performance may vary. NG‑3 should only be deployed with proper clinical oversight and in compliance with all applicable healthcare regulations including HIPAA.