🎓 Research Scientist & Machine Learning Engineer
📍 Based in France
🌐 emanuelaboros.github.io
📚 Google Scholar ·
ORCID
I work at the intersection of natural language processing, machine learning, and document analysis.
My research focuses on building and evaluating models that remain useful when the data is not clean, balanced, contemporary, or well-resourced. I am especially interested in multilingual and historical documents affected by noise, domain shift, temporal variation, and incomplete supervision. My research is guided by three connected questions:
How can NLP and language models accurately extract and reason over information when documents are noisy, multilingual, historical, or low-resource? How should we evaluate these systems so that metrics reflect semantic usefulness, robustness, calibration, and downstream impact—not only performance on clean benchmarks? How can models adapt across languages, time periods, domains, and document modalities while remaining grounded, interpretable, and reproducible?
I explore these questions through dataset construction, controlled-noise experiments, task-aware evaluation, error analysis, model adaptation, and information extraction.
Building multilingual systems for named entity recognition, entity linking, relation extraction, and event extraction over noisy and heterogeneous documents.
Designing task-aware metrics, benchmarks, calibration methods, and error analyses that measure semantic usefulness and downstream impact—not only surface similarity.
Studying how language models behave under OCR/HTR/ASR noise, temporal variation, domain shift, multilinguality, and limited supervision.
- multilingual information-extraction systems;
- historical language-model adaptation methods;
- reproducible notebooks, teaching resources, and open research tools;
- evaluation frameworks, benchmarks, and task-aware metrics;
- GPU-based training and large-scale experimentation workflows;
- LLM and OCR/HTR post-correction experiments;
- datasets, annotation resources, and synthetic-data pipelines.




