openmed.analyze_text is the top-level orchestrator that most users start with. It validates input, spins up a
token-classification pipeline, segments sentences, and normalizes the output so you can copy dict/JSON/HTML/CSV payloads
straight into downstream systems.
from openmed import analyze_text
result = analyze_text(
text="Patient started on imatinib for chronic myeloid leukemia.",
model_name="disease_detection_superclinical",
aggregation_strategy="simple",
output_format="dict",
include_confidence=True,
confidence_threshold=0.55,
group_entities=True,
metadata={"source": "clinic-note-42"},
)
print(result.model)
print(result.entities[:3])
payload = result.to_dict()model_name: registry alias, full Hugging Face id, or local model directory. Useopenmed.list_models()if you need auto-discovery.model_id: alias formodel_name, supported for API-style callers that use model-id terminology.aggregation_strategy: forwarded to the HF pipeline.simple(default) yields grouped tokens;Nonekeeps raw tokens.output_format:"dict"(default, returnsAnalyzeResult),"json","html", or"csv".include_confidence&confidence_threshold: control the final payload; defaults keep all scores.group_entities: merge adjacent spans of the same label after formatting.formatter_kwargs: forwarded toopenmed.processing.format_predictions.assert_context: opt in to deterministic clinical assertion labels. Each entity receivesnegation,uncertainty,experiencer, andtemporalityunderentity.metadata["clinical_context"].- Sentence options (
sentence_detection,sentence_language,sentence_clean,sentence_segmenter,sentence_backend) wrap the sentence engine so each prediction carries the sentence span; disable them if latency matters more than helper metadata.
The context stage is disabled by default. Enable it for clinical entities that will flow into grounding, FHIR export, or problem-list review:
result = analyze_text(
"No evidence of pneumonia.",
assert_context=True,
)
print(result.entities[0].metadata["clinical_context"])The default sentence_backend="auto" path is unchanged: OpenMed uses its built-in Indic and Chinese segmenters where
appropriate and pySBD elsewhere. YASBD is neither installed nor imported by a core OpenMed installation.
Install the experimental backend explicitly when you want to benchmark it on your own workload:
pip install "openmed[yasbd]"Then select it for either the low-level sentence API or analyze_text:
from openmed import analyze_text
from openmed.processing.sentences import segment_text
spans = segment_text(
"Patient is stable. Follow up tomorrow.",
language="en",
backend="yasbd",
)
result = analyze_text(
"Patient is stable. Follow up tomorrow.",
sentence_backend="yasbd",
)The adapter preserves OpenMed's exact, contiguous source offsets and assigns inter-sentence whitespace to the preceding
span, matching the existing span contract. YASBD remains an explicit opt-in because sentence boundaries can differ
between engines; validate representative clinical and multilingual inputs before adopting it in production. If the extra
is missing, explicitly selecting "yasbd" raises an installation error instead of silently changing behavior.
result = analyze_text(
text=long_report,
model_name="pharma_detection_superclinical",
aggregation_strategy=None, # work with raw tokens
max_length=512, # forwarded to HF pipeline
truncation=True, # enforce length (default)
sentence_detection=False, # skip sentence detection to save ~2ms per note
sentence_backend="auto", # "auto" (default) or "yasbd" (experimental, faster)
)When you need full-control over tokenizer behaviour:
- Pass
max_length/truncationviapipeline_kwargs. If you skip truncation, the helper sets the tokenizer max length to unlimited (0) so HF pipelines accept longer inputs. - Provide
batch_sizeornum_workersinpipeline_kwargsand they will be forwarded to the pipeline call but not to the constructor. - Enable medical token remapping with
OpenMedConfig(use_medical_tokenizer=True)to group outputs onto clinical-friendly tokens without changing the model tokenizer.
Pass an existing model directory to model_name or model_id when the model files are already present on disk:
import os
from openmed import OpenMedConfig, analyze_text
local_path = os.path.abspath("./models/OpenMed-NER-DiseaseDetect-SuperClinical-434M")
config = OpenMedConfig(device="cpu")
result = analyze_text(
"Patient presents with chronic myeloid leukemia and Type 2 diabetes.",
model_id=local_path,
config=config,
)
for entity in result.entities:
print(entity.text, entity.label)
legacy_payload = result.to_dict()
print(legacy_payload["model_name"])When the identifier points to an existing local path, OpenMed asks Transformers to load with local_files_only=True by
default. That keeps air-gapped deployments from validating or downloading the model from the Hugging Face Hub. If any
required tokenizer, config, or weight file is missing, loading fails locally with the underlying Transformers error.
analyze_text is optimized for single inputs. For batch jobs, keep a ModelLoader instance around and reuse its
pipelines:
from openmed import ModelLoader, format_predictions
loader = ModelLoader()
pipeline = loader.create_pipeline("disease_detection_superclinical")
for note in notes:
raw = pipeline(note, batch_size=16)
formatted = format_predictions(raw, note, model_name="Disease Detection")
print(formatted.entities[:3])See ModelLoader & Pipelines for details on caching, GPU selection, and tokenizer reuse.
html = analyze_text(
text,
model_name="oncology_detection_superclinical",
output_format="html",
formatter_kwargs={
"html_class": "openmed-highlights",
"tag_colors": {"CANCER": "#d97706"},
},
)
csv_rows = analyze_text(
text,
model_name="pharma_detection_superclinical",
output_format="csv",
)The HTML formatter emits a ready-to-embed snippet for dashboards; CSV mode writes row strings (header + body). Both respect
confidence_threshold and group_entities.
Behind the scenes analyze_text calls:
validate_input— trims whitespace and enforces max lengths.validate_model_name— normalizes registry aliases.- Sentence detection (
openmed.processing.sentences) — optional segmentation with language hints and a selectable backend. OutputFormatter— see Advanced NER & Output Formatting for available kwargs.
If you need custom validation or logging, inject your own OpenMedConfig or reuse a configured ModelLoader.