OntoBook: Using Ontology-Generated Textbooks to Improve Medical Encoder Pretraining
Researchers present OntoBook, a method that transforms medical ontology structures into synthetic textbook-style prose using large language models, and uses this data to pretrain a French medical encoder (ModernCamemBERT). The approach leads to significant improvements on three French medical coding benchmarks, including up to +8.0 micro-F1 on Distemist-FR. The team also releases 1.3 million generated textbooks and pretrained model checkpoints.
Why it matters: This work offers a scalable way to inject structured medical knowledge into language models, potentially advancing clinical NLP in low-resource languages.
Full story at: arXiv AI/ML ↗