From Tables to Knowledge: Extracting Pharmacokinetic Data from Literature Tables with Natural Language Processing
Abstract
Efficient extraction and integration of pharmacokinetic (PK) data from scientific literature is critical for informed decision-making in drug development. Prior knowledge of PK parameters, particularly from similar compounds, supports first-in-human dosing, parameter estimation, and compound screening, ultimately helping to reduce attrition in clinical trials. While recent natural language processing (NLP) efforts have focused on extracting PK data from unstructured text, these approaches often overlook more comprehensive PK information and essential contextual metadata, which are usually reported in tables. Despite the prevalence and value of these tables, no previous work has systematically addressed the automated extraction of PK data from them. This thesis presents a novel NLP pipeline for identifying, extracting, and structuring PK data from scientific tables. The work addresses a key gap by targeting tables as a rich and underutilised source of PK information. The thesis is structured around four main components. First, a classification system combining supervised learning and prompt-based approaches is developed to retrieve PK-relevant tables from full-text biomedical articles. Second, heuristic and neural named entity recognition approaches are designed to extract PK parameters and associated metadata from table cells, including dose, species, study population, route of administration, units, and other contextual qualifiers. Third, an entity linking pipeline, including rule-based and zero-shot methods, is developed to normalise extracted data to a standardised PK ontology. Finally, the full pipeline is utilised to construct a large-scale PK database from PubMed Open Access articles. The database is evaluated through systematic sampling and manual quality assessment, and proof-of-concept analyses demonstrate how the extracted data can be used to characterise literature-wide reporting trends and explore comparative pharmacological questions. The results of this thesis demonstrate that automated PK table mining is both feasible and scalable, significantly accelerating the curation of high-quality datasets for pharmacometrics modelling. This work presents new open-source annotated corpora, domain-specific NLP methodologies, and practical tools for structuring PK literature, thereby opening the door to scalable, data-driven approaches in early drug development.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.