The main players, opportunities for the cardiologist
Which tool do you pick when? Per ChatGPT (OpenAI), Claude (Anthropic) and Gemini (Google) what makes them strong and where they fall short. Plus Mistral, the French/EU player bringing sovereignty within the EU, and DeepSeek (CN) with the sovereignty question. And why the pace is not linear but exponential.
What you'll learn: strengths and pitfalls per tool, a choice matrix per cardio task, and how not to get stuck with one vendor.
Lesson 3.1: ChatGPT: the generalist
OpenAI · since Nov 2022 · 700M+ weekly users · GPT-5.6 in ChatGPT (2026); GPT-6 Astra is rolling out.
- GPT-5.6: one family, two speeds. Router between Instant (fast) and Thinking (deep reasoning). Significantly fewer hallucinations than GPT-4o.
- Native image generation. gpt-image-2 (2026), strong text-in-image for educational material and infographics.
- Advanced Voice mode. Real-time conversation; discuss a case on the go without typing (note: no patient data).
- Code Interpreter. Python in the chat. Upload anonymised CSV → have it analysed and visualised.
- Agent / Operator. Executes tasks in a browser, books, fills in forms, scrapes data.
Caveats
- Fast release cycle: quality varies per feature and date.
- Image-gen often refuses medically recognisable content (copyright / person protection).
- Voice and agent don't (yet) run inside the hospital perimeter, no patient data.
- The “memory” feature remembers info across chats. Useful or worrying, depending on what you share.
Questions for lesson 3.1
1. For which task is ChatGPT (GPT-5.6 + image-gen) usually the first choice among the current players?
Lesson 3.2: Claude: the text worker
Anthropic · since March 2023 · Opus 5 (July 2026); Fable 5.1 and Mythos 5.1 (September 2026).
- Excel, PowerPoint & Word add-ins (M365, since late 2025): edits sheets, formulas, slides and documents directly in the file; marks every change.
- Artifacts. Interactive HTML/React/SVG live next to the chat, build a widget, dashboard or poster while you talk.
- Long context. 1M tokens on current Claude models (Fable 5.1 / Mythos 5.1). Older tiers: 200K. A complete ESC guideline fits without cutting.
- Coding leader. Fable 5.1 and Opus 5: current coding and agent tiers (2026). Strong for scripts, data pipelines, refactoring.
- MCP integrations. 6,000+ connectors: Drive, Slack, GitHub, Jira.
The ethical positioning, what it does and doesn't mean
Anthropic publicly positions itself emphatically around AI safety and responsible use (including themes like Constitutional AI: the model is given explicit principles to follow). In practice: Claude is often more cautious than ChatGPT with direct medical advice. That is deliberate design. It means the quality of safety discussions is often a bit higher, but it is no carte blanche to share patient data. Claude hallucinates too, Claude is also US-based, and you remain responsible.
Caveats
- No native image generation, no video, text (and image) in, text out.
- No voice mode like ChatGPT.
- Web interface is cleaner but more feature-poor than ChatGPT for consumers.
Questions for lesson 3.2
2. What's a typical sweet spot for Claude (Fable 5.1)?
3. “Constitutional AI”, Anthropic's much-mentioned label, is essentially:
Lesson 3.3: Gemini: the Workspace dweller
Google DeepMind · since Dec 2023 (formerly Bard) · Gemini 3.1 Pro and Gemini 3.8 Flash (2026).
- Workspace integration. Lives in Gmail, Docs, Sheets, Slides, Drive, Meet, doesn't ask for context, already has access.
- Personal Intelligence. Knows (with permission) your mail, calendar, Drive; prepares meetings, summarises threads.
- Deep Research (Max). Autonomous research agent. Searches web + Workspace, delivers a report with citations.
- Native video & multimodal. Processes video directly, not just frames, suitable for echo clips or ward recordings (with safe framework).
- Veo / Nano Banana / Lyria. Integrated image, video and music generation. For educational or presentation material.
Caveats
- Workspace connection = your Google data is read by Gemini. Check your organisation settings.
- Standalone chat: experienced by many as weaker in prose than ChatGPT / Claude.
- Product line with many names: Gemini app, Google AI Pro, Ultra, Workspace, NotebookLM, confusing.
- EU rollout of Personal Intelligence lags due to data-residency legislation.
Lesson 3.4: Mistral / Le Chat: the European player
Mistral AI · since 2023 · Paris (France, EU) · Le Chat (chat) · Large 2, Codestral and open-weights models.
Mistral was founded in 2023 by three former DeepMind/Meta researchers and quickly grew into the European counterweight to OpenAI/Anthropic/Google. For the cardiologist dealing with “data within Europe”, this is often the first name to appear on an EU hospital's vendor list.
- EU-based. Head office in Paris; primary jurisdiction within the EU, so GDPR as a starting point rather than a later point of contention.
- Data residency manageable. For business subscriptions and partner deployments, EU hosting can be arranged, an important argument for DPIA and NEN-7510.
- Open-weights variants. Models like Mistral Small and Mixtral are available as open weights; hospitals can in principle consider self-hosted or partner-hosted deployments.
- Le Chat. Web/app interface with features comparable to ChatGPT/Claude: search, image, code, documents.
Caveats
- Smaller ecosystem. No M365 add-in like Claude, no Workspace integration like Gemini. At most loose plug-ins.
- Clinical evaluation thinner. Countless published medical evaluations exist for ChatGPT and Claude (USMLE, ESC cases, radiology); for Mistral that field is still underrepresented.
- “EU” ≠ all data stays in the EU. Check per product, subscription and sub-processor what the actual processing location is. Mistral is also offered via Azure, then Microsoft is the processor, not Mistral itself.
- Competition stays sharp. Mistral's release pace is high but R&D budgets smaller than the American top, differences on benchmarks can shift each quarter.
The practical rule for cardiologists
For departments where data residency and EU jurisdiction weigh heavily: ask ICT/CIO whether an organisation-approved Mistral route (Le Chat Enterprise, or Mistral via an EU Azure region) is available. For pure quality on letters, guidelines and reasoning, ChatGPT and Claude are often still stronger today, but the difference is shrinking, and the GDPR argument is hard to wave away.
Questions for lesson 3.4
1. What is the most important practical argument for a European cardiology department to put Mistral on the short list?
2. You want to edit a draft discharge letter directly in Word (track changes on), in an EHR environment running on Microsoft 365. Which tool fits best today?
Lesson 3.5: DeepSeek and the sovereignty question
DeepSeek (深度求索) · China, Hangzhou (Zhejiang).
DeepSeek-V3 and R1 received much attention in 2025 because the models include open weights and benchmark comparably to Western top models, at a fraction of the reported training costs. For the cardiologist the discussion is less “is it good?” and more “what is the legal and organisational context?”
- Chinese company, headquartered in Hangzhou (Zhejiang province).
- Processing via the official chat/API runs through Chinese servers, with Chinese legal context.
- The models are partly open weights: local and EU-hosted deployments exist via third parties, those have different terms.
- Same technical risks as other LLMs: hallucination, bias, “too confident” text.
The practical rule
For healthcare institutions in Europe: typically no patient data in the public DeepSeek chat. For research or comparison with anonymised cases it can be interesting, but first check with ICT/CIO whether any form of EU-hosted or self-hosted DeepSeek is available.
Questions for lesson 3.5
4. What is most important practically for a European cardiologist regarding DeepSeek?
Lesson 3.6: Which tool when? Cardio choice matrix
Rules of thumb. For specific tasks, differences can play out differently, test for yourself with your use cases.
Free versus paid: the matrix below is about strength per task, not about what the free tier can handle. Long guidelines, Projects and the heaviest reasoning mode often sit behind Plus, Pro or Google AI Pro. A short differential or a paragraph can be done in the free chat on all three. What forums and review sites crown this month will shift; verify in your own workflow, and read the vendor page for current free limits.
| Task | First choice | Why |
|---|---|---|
| Edit discharge letter or report | All (incl. Mistral) | All produce good first drafts. Always verify yourself. |
| Strict EU data residency / GDPR-first department policy | Mistral | Only one of the big four with main establishment within the EU; easier to substantiate for DPIA and NEN-7510. |
| Analyse a whole guideline / file at once | Claude | 200K–1M tokens of context: long documents without cutting. |
| Working IN Excel, PowerPoint, Word | Claude | M365 add-in (since late 2025): edits sheets, formulas and slides directly in the file. |
| Working IN Gmail, Docs, Sheets (Google Workspace) | Gemini | Only one with deep Workspace integration. |
| Python / R script or data analysis | Claude or ChatGPT | Claude leads on SWE-bench; ChatGPT has Code Interpreter (Python in the chat, sandbox). |
| Search & summarise literature with citations | Gemini | Deep Research autonomously consults 50+ sources with citations. |
| Create image or poster (education) | ChatGPT | gpt-image-2 (2026): best text rendering in image. |
| Interactive HTML tool or dashboard prototype | Claude | Artifacts: live preview and shareable as an app. |
| Practise voice chat / discuss a case without typing | ChatGPT | Advanced Voice Mode (no patient data). |
Cardio-specific: a comparative study (Hearts, 2025, Di Eusanio et al.) had ChatGPT score slightly higher than Claude/Gemini on clinical prompts, differences small. Mistral was not included in that study; first own tests therefore call for cautious conclusions. More important than tool choice: whether you verify.
Questions for lesson 3.6
5. You want to clean up an MDT PowerPoint file, rewrite one slide and keep the form in track changes. Which tool do you pick today?
6. What is a sensible selection attitude for a cardiology department wanting to work with LLMs in 2026?
Live exercise: compare two tools on the same question
Ask exactly the same question to ChatGPT and to Claude and carefully compare the answers (no patient data).
Note: (1) does CHA2DS2-VASc come out the same in both tools? (2) which gives a guideline citation you can't verify? (3) which feels clinically more “mainstream”? No patient data; the above is a fictional case.
Lesson 3.7: The pace: not linear, exponential
Goal: understand why you must adjust your expectations every 6–12 months.
Capabilities that were science fiction in 2022 are today's toolbar. One indicative graph from medical evaluations:
| Year | Model | USMLE / MedQA-like benchmark |
|---|---|---|
| 2020 | GPT-3 | Below pass rate |
| 2023 | GPT-4 | Above USMLE pass rate |
| 2025–2026 | GPT-5 / Claude Opus 4.7 | 95%+ on academic exams |
| 2026 | GPT-5.6 / Claude Fable 5.1 | Frontier; context windows around 1M tokens |
Source: illustrative, based on public evaluations (Singhal et al. 2023; Kung et al. 2023; more recent comparative studies 2025–2026). Benchmarks ≠ clinical performance. Telling about direction; little about whether it works in your specific department.
What to take from this
Expect that what is “not quite usable” today can often be usable 12 months from now. At the same time: newer models also become subtler in their errors, a “not quite” error in a guideline citation sometimes stands out less than a gross one. Verify what matters remains, even with better models, the central working attitude.
Take-home from module 3
The players
ChatGPT = generalist + image + voice. Claude = long context + M365 + ethical framing. Gemini = Workspace + Deep Research + native video. Mistral = the EU route, strong for data residency / GDPR. DeepSeek = open-weights from CN with sovereignty discussion.
Tool ≠ goal
Choose per task. Organisational policy and DPIA come before benchmarks. DeepSeek not as a public chat for patient data; Mistral can make the difference in EU contexts with “data within Europe” requirements.
Test every 6 months
What didn't work last year often works now. What seems “all good” now could soon be better via another model.