AI tool may help decision-making in therapeutic plasma exchange

RAG-based AI outperformed standard models in simulated therapeutic plasma exchange consultations.

New research suggests that an artificial intelligence-based consultation system may help doctors make more accurate and consistent decisions about therapeutic plasma exchange (TPE), a blood-filtering treatment used in some immune-related, neurologic, kidney, and blood disorders.

The system, called retrieval-augmented generation for therapeutic plasma exchange (RAG-TPE), uses an approach in which an artificial intelligence tool searches trusted medical guidance before generating an answer, rather than relying solely on the language model’s built-in knowledge.

Although the study did not include fetal and neonatal alloimmune thrombocytopenia, the findings may be relevant to transfusion medicine more broadly, where antibody-related conditions are an important area of care. However, the system was tested in a simulated consultation setting, not during real-time patient care.

Therapeutic plasma exchange (TPE) is a procedure that removes and replaces plasma, the liquid part of blood. It may be used when harmful antibodies, immune proteins, or toxins need to be removed from the bloodstream.

Read more about FNAIT treatment and care

In the study, researchers developed the RAG-TPE system using the 2023 American Society for Apheresis (ASFA) guidelines, which help clinicians decide when TPE is appropriate. The system was also designed to consider insurance criteria, plasma volume calculations, and replacement fluid recommendations.

Researchers tested the system using 30 real-world consultation cases that had been stripped of identifying patient information and rewritten as standardized clinical questions. They compared 6 RAG-based artificial intelligence models with 3 standard large language model configurations that did not use retrieval. Each case was answered multiple times, producing 1,350 total outputs.

The outputs were tested for accuracy across 6 decision elements: diagnosis, ASFA category, ASFA grade, insurance applicability, plasma volume, and replacement fluid. RAG-based models were more accurate overall than standard models, with an average accuracy of 88.6% compared with 61.3% for non-RAG models.

The largest improvements were seen in plasma volume calculations and ASFA classification, 2 areas where incorrect or inconsistent answers could affect clinical decision-making. The best-performing configuration, RAG GPT-4.1-mini, reached 95.5% mean accuracy while maintaining relatively fast response times.

RAG-based models also gave more consistent answers when tested repeatedly. This finding is important because clinical decision support tools need to provide stable recommendations, rather than different answers each time the same medical question is asked.

The authors stated future research should test the system in hospital workflows, improve the user interface, connect it with hospital information systems and explore local large language models to support clinical usefulness, safety and scalability.