DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowFall ResetAmazon USFall reset deals: check better picks before checkoutAmazon US: today's deals, useful picks and quick comparisons.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Blog

Writer’s Palmyra AI Models Posted Striking Healthcare and Finance Scores. What Do They Prove?

By TheFinanceBase Team7 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Writer’s Palmyra-Med-70B and Palmyra-Fin-70B delivered attention-grabbing results on healthcare and finance tests—but those results are company-reported benchmarks, not proof that the models can safely diagnose patients or manage investments. Announced on July 31, 2024, the models were built for specialized tasks such as medical-document analysis and financial research. Writer said Palmyra Med averaged 85.9% across its medical benchmarks and Palmyra Fin scored 73% on the multiple-choice portion of a CFA Level III sample exam. Those numbers suggest domain expertise worth evaluating, not autonomous professional judgment.

What Writer released

Writer announced two roughly 70-billion-parameter-class language models: Palmyra-Med-70B for healthcare and Palmyra-Fin-70B for financial services. The “70B” label indicates model scale; it is not, by itself, evidence of quality. Writer positioned Med for work involving clinical notes, electronic health records, discharge summaries, biomedical research, medical coding, and clinical-trial information. Fin was aimed at financial-document analysis, investment research, risk assessment, reporting, and related workflows.

The models were introduced through Writer’s platform and API, with launch announcements also describing availability through NVIDIA, Baseten, and Hugging Face. Writer called them open models, but that does not necessarily mean unrestricted commercial use: its announcement directed commercial users to Writer for licensing information. Writer’s launch announcement describes the original release and distribution channels.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Specialized models are appealing because healthcare and finance rely on dense terminology, long documents, and structured processes, while errors can carry unusually high costs. A model tuned for these domains may understand their language and recurring document patterns better than a general-purpose model on a particular task. That possibility still needs to be demonstrated on an organization’s own data and workflow.

What the healthcare scores show—and what they don’t

Writer reported that Palmyra Med averaged 85.9% across its medical benchmarks, and scored 80% on PubMedQA, a biomedical question-answering benchmark. Writer also compared its result with Med-PaLM 2, citing an approximately 84% score and saying that Palmyra Med reached its result zero-shot while the comparison used examples or multiple attempts. These are results as described by the vendor, not an independent ranking. The details needed to fully assess the comparison—including test composition, exact prompts, repeated runs, confidence intervals, and error analysis—are not established by the headline scores alone. See Writer’s engineering account of its evaluations.

A benchmark can show that a model answers a defined set of questions well under defined conditions. It does not show that the model is safe across hospitals, specialties, patient populations, languages, or clinical settings. An average may also hide weaker performance on particular tasks. These scores do not establish regulatory approval, diagnostic safety, or how well the system handles incomplete or contradictory records, misleading inputs, or local clinical policies. Clinical deployment calls for prospective validation, strong controls, and qualified human oversight.

That distinction matters even for seemingly routine uses. A summary that drops a contraindication or an important qualifier can mislead a clinician; a medical-code suggestion can be wrong because similar terms have different meanings. A model can assist with finding and organizing information, but a fluent answer is not evidence that it checked the underlying record correctly.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the finance scores show—and what they don’t

Writer said Palmyra Fin scored 73% on the multiple-choice section of a CFA Level III sample examination, compared with an approximately 60% average for human test takers over an 11-year period. Writer described the experiment as an ad hoc, zero-shot test using a sample exam in PDF form, with questions presented one at a time. This was not the complete official examination, and the comparison does not establish equivalent testing conditions or broad professional competence.

Writer also said Palmyra Fin outperformed GPT-4o, Claude 3.5 Sonnet, and Mixtral on its long-fin-eval benchmark. Because Writer described that evaluation as internally created, its construction, scoring, contamination controls, and independent verification are important considerations. The company’s published comparisons do not, on their own, establish a neutral and reproducible ranking. The methodology and claims are in Writer’s testing write-up.

A strong exam result is not evidence of profitable investment decisions. Real financial work involves current data, uncertainty, market impact, client suitability, legal duties, and accountability. The exam score does not prove that the model can forecast markets, allocate assets, detect fraud reliably, or give appropriate investment advice. Those are proposed application areas, not outcomes demonstrated by a sample test.

How to read claims that a specialist model “beats” a general one

“Outperforms GPT-4” is incomplete unless the comparison specifies exactly what was tested. A buyer should ask which model versions were used, whether every model received the same prompt, and whether tools, retrieval, external data, images, and token limits were comparable. It also matters whether the task was zero-shot or few-shot, whether the benchmark was public or private, whether the run was repeated, and what the score measured: answer selection, factuality, citation quality, calibration, or completion of a real workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Public exam questions and datasets also raise a possible training-data contamination issue. That possibility does not invalidate a result, but it limits what can be inferred unless test-set controls are explained. A score on a benchmark is best treated as a lead for further evaluation—not a substitute for one.

Where a model may be useful

The most defensible starting points are bounded information-processing tasks where a person can check the output against a source:

  • Lower-risk starting uses: summarize internal documents, extract fields, classify records, locate passages, convert text into structured data, or draft internal reports for review.
  • Uses that call for robust review: suggest medical codes, summarize drug-interaction information, analyze clinical-trial documents, draft risk or regulatory reports, summarize investment research, or triage fraud alerts.
  • High-consequence decisions: diagnosis and treatment, autonomous patient communication, credit or insurance decisions, final investment recommendations, trading, fraud accusations, and regulatory determinations should not be delegated to a language model acting alone.

The practical dividing line is assistance versus autonomy. For an assistive workflow, the model can prepare or organize information while a qualified person remains responsible for checking it. An autonomous system that makes or communicates consequential decisions needs a much higher standard of validation, governance, and accountability.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What has changed since the 2024 launch

The “new models” label is historical: Med and Fin were announced in July 2024. Writer’s current catalog also includes newer models such as Palmyra X5, which Writer describes as a general-purpose, agent-oriented model with a one-million-token context window. The current catalog lists Palmyra Med at 32K tokens and Palmyra Fin at 128K. A larger context window can accommodate more input, but does not guarantee the model will find, weigh, or correctly use every relevant passage. Writer’s current model catalog lists its lineup and specifications; the X5 announcement covers that newer model.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

As listed on Writer’s site on August 18, 2026, Palmyra Fin cost $5 per million input tokens and $12 per million output tokens. Writer listed X5 at $0.60 per million input tokens and $6 per million output tokens in its announcement. Prices and availability can change, and the X5 price is not a like-for-like comparison proving lower total cost for a given task. Check Palmyra Fin’s current page before budgeting. Total cost also includes retrieval, integration, hosting where relevant, monitoring, human review, and compliance work.

Writer describes its platform as offering tools such as retrieval, guardrails, structured output, and tool calling, and says it does not use customer-shared data to train or modify its models. These are vendor statements and platform-level capabilities, not guarantees about every product tier, contract, region, or implementation. Confirm data processing, retention, deletion, access controls, logging, and deployment terms in the applicable agreement. A platform’s controls do not automatically make a particular use compliant with healthcare privacy or financial-services obligations.

How an enterprise should evaluate Palmyra

Before choosing a specialist model, test it against representative, permitted examples from the actual workflow and compare it with relevant alternatives under the same conditions. Separate results for extraction, summarization, classification, and reasoning; report false positives and false negatives rather than only an overall average. Check whether outputs cite source passages, preserve qualifiers, flag missing information, and admit uncertainty. Include difficult cases: stale data, contradictory records, scanned tables and footnotes, ambiguous terminology, adversarial instructions hidden in documents, and incomplete files.

Then evaluate operational fit: where data is processed, whether it is retained or used for training, what audit logs are available, whether administrators can require review, and whether prompts, retrieved sources, outputs, and model versions can be traced. Check integration with the relevant records and databases, freshness of retrieved data, and controls against unauthorized actions. Measure latency and throughput on realistic documents, not just token costs or context-window size. If a model suggests codes, risk flags, or other consequential outputs, define who verifies them and what happens when the model is wrong.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Specialization is a trade-off, not a universal advantage. Palmyra may be a good fit if it measurably improves a defined medical or finance workflow. A general-purpose model may be more useful for work spanning many unrelated domains or where its surrounding tools better meet the need. A self-hosted model may offer more deployment control but require greater engineering and operational effort. The appropriate choice depends on task-specific accuracy, governance, integration, and total cost—not a single benchmark chart.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Written by TheFinanceBase Team

The Team behind TheFinanceBase.

Add your note

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.