Novel document-level measures and their performance for active learning uncertainty sampling: a use case for automatic CDSS ontology curation
- Open access
- 1 cites
Novel document-level uncertainty aggregation strategies significantly improve active learning performance in ontology curation, with KPSum showing consistent gains over random sampling.
- Why it matters: Effective and transparent document selection is crucial for automating ontology development and reducing annotation effort, addressing gaps in current active learning methods focused on token-level uncertainty.
- What they did: Using a BILSTM-CRF model trained on over 2900 PubMed abstracts, four new document-level uncertainty strategies—KPSum, KPAvg, DOCSum, and DOCAvg—were evaluated against random sampling, measuring recall and F1 scores.
- The result: KPSum (actual order) consistently enhanced early-cycle recall and F1, demonstrating potential to automate ontology maintenance by prioritizing impactful documents, though further improvements are needed for real-world deployment.