Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “identifiability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

NCAPH drives breast cancer progression and identifies a gene signature that predicts luminal a tumour recurrence

Luminal A tumours generally have a favourable prognosis but possess the highest 10-year recurrence risk among breast cancers. Additionally, a quarter of the recurrence cases occur within 5 years post-diagnosis. Identifying such patients is crucial as long-term relapsers could benefit from extended hormone therapy, while early relapsers might require more aggressive treatment. We conducted a study to explore non-structural chromosome maintenance condensin I complex subunit H’s (NCAPH) role in luminal A breast cancer pathogenesis, both in vitro and in vivo, aiming to identify an intratumoural gene expression signature, with a focus on elevated NCAPH levels, as a potential marker for unfavourable progression. Our analysis included transgenic mouse models overexpressing NCAPH and a genetically diverse mouse cohort generated by backcrossing. A least absolute shrinkage and selection operator (LASSO) multivariate regression analysis was performed on transcripts associated with elevated intratumoural NCAPH levels. We found that NCAPH contributes to adverse luminal A breast cancer progression. The intratumoural gene expression signature associated with elevated NCAPH levels emerged as a potential risk identifier. Transgenic mice overexpressing NCAPH developed breast tumours with extended latency, and in Mouse Mammary Tumor Virus (MMTV)-NCAPH ErbB2 double-transgenic mice, luminal tumours showed increased aggressiveness. High intratumoural Ncaph levels correlated with worse breast cancer outcome and subpar chemotherapy response. A 10-gene risk score, termed Gene Signature for Luminal A 10 (GSLA10), was derived from the LASSO analysis, correlating with adverse luminal A breast cancer progression. The GSLA10 signature outperformed the Oncotype DX signature in discerning tumours with unfavourable outcomes, previously categorised as luminal A by Prediction Analysis of Microarray 50 (PAM50) across three independent human cohorts. This new signature holds promise for identifying luminal A tumour patients with adverse prognosis, aiding in the development of personalised treatment strategies to significantly improve patient outcomes.

60 APPLIED LIFE SCIENCES↗

Using Neural Networks to Identify Mixture Components in Hyperspectral Reflectance Data

Neural networks have been employed to identify materials of interest from hyperspectral data (generally imagery) based on their unique spectral signatures. This approach assumes that there is a single material that is standing out from the rest of the spectrum to be identified. However, pixels often contain more than one material, or a material of interest may itself be a mixture of multiple materials. Neural networks are only as good as the data used to train them, and it takes a great deal of work in the laboratory to identify, make, and measure all potential mixtures of interest. Thus, researchers often calculate synthetic spectra using algorithms with varying degrees of fidelity to the physics that govern the interactions between light and multiple materials. In this work, we have (1) adapted a neural network designed to identify mixture components from Raman spectroscopy to work with visible to near‐infrared reflectance data and (2) tested three common mixture algorithms to determine the most accurate and least computationally expensive method to build synthetic training datasets. With our initial test dataset, we have achieved accuracies of > 90% and found that the synthetic training dataset produced using the Hapke mixture model provides the best results.

99 GENERAL AND MISCELLANEOUS↗

A proposed criteria to identify wind turbine drivetrain bearing loads that induce roller slip based white-etching cracks

In this article, the type of roller slip behavior that may result in the formation of white-etching cracks (WECs) in wind turbine gearbox bearings is identified. A new hypothesis based on the inner raceway normal contact load magnitude at the time of roller slip is proposed as the probable cause of WECs. For this purpose, the maximum normal contact loads are identified when roller slip occurs in high-speed shaft bearings at different mean wind speeds. Subsequently, the annual probability of occurrence of the maximum normal loads are obtained. The probability of maximum load under slip exceeding a limit probability is hypothesized as a probable cause for WEC. In order to apply the proposed hypothesis, two different wind turbines high-speed shaft bearings are used: the cylindrical roller bearing of the General Electric 1.5 SLE turbine and the tapered roller bearing of Vestas V52 turbine. Both the chosen bearings are on the generator side of the high-speed shaft. For both turbines, measurement data together with analytical models are used for identifying the slip and the maximum normal contact loads. We propose forecasting the probability of exceedance of a threshold maximum normal contact load level during slip to identify the possibility for inducing WECs.

17 WIND ENERGY↗

Integrated Transcriptomic and Proteomic Analysis Identifies Plasma Biomarkers of Hepatocellular Failure in Alcohol-Associated Hepatitis

Alcohol-associated hepatitis (AH) is a form of liver failure with high short-term mortality. Recent results have shown that HNF4a defective function and systemic inflammation are major disease drivers of AH. Plasma biomarkers of hepatocyte function could be useful for diagnostic and prognostic purposes. Herein an integrative analysis of hepatic RNAseq and liquid chromatography-tandem mass spectrometry (LC-MS/MS) was performed to identify plasma protein signatures for mild and severe AH patients. Alcohol-related liver disease cirrhosis (ALD)(AC), non-alcoholic fatty liver disease (NALFD), and healthy subjects (HC) were used as comparator groups. Identified proteins primarily involved in hepatocellular function were decreased in AH patients which included hepatokines, clotting factors, complement cascade components, and hepatocyte growth activators. A protein signature of AH disease severity was identified including thrombin (THRB), hepatocyte growth factor alpha (HGFA), clusterin (CLUS), human serum factor H-related protein (FHR1) and kallistatin (KAIN), which exhibited large abundance shifts between severe and non-severe AH. The combination of THRB and HGFA discriminated between severe and non-severe AH with high sensitivity and specificity. These findings were correlated with the liver expression of genes encoding secreted proteins in a similar cohort, finding a highly consistent plasma protein signature reflecting HNF4A and HNF1A functions. This unbiased proteomic-transcriptome analysis identified plasma protein signatures and pathways associated with disease severity, reflecting HNF4A/1A activity useful for diagnostic assessment in AH.

60 APPLIED LIFE SCIENCES↗

Multiomic Network Analysis Identifies Dysregulated Neurobiological Pathways in Opioid Addiction

BACKGROUND: Opioid addiction is a worldwide public health crisis. In the United States, for example, opioids cause more drug overdose deaths than any other substance. However, opioid addiction treatments have limited efficacy, meaning that additional treatments are needed. METHODS: To help address this problem, we used network-based machine learning techniques to integrate results from genome-wide association studies of opioid use disorder and problematic prescription opioid misuse with transcriptomic, proteomic, and epigenetic data from the dorsolateral prefrontal cortex of people who died of opioid overdose and control individuals. RESULTS: Here we identified 211 highly interrelated genes identified by genome-wide association studies or dysregulation in the dorsolateral prefrontal cortex of people who died of opioid overdose that implicated the Akt, BDNF (brain-derived neurotrophic factor), and ERK (extracellular signal-regulated kinase) pathways, identifying 414 drugs targeting 48 of these opioid addiction–associated genes. Some of the identified drugs are approved to treat other substance use disorders or depression. CONCLUSIONS: Our synthesis of multiomics using a systems biology approach revealed key gene targets that could contribute to drug repurposing, genetics-informed addiction treatment, and future discovery.

60 APPLIED LIFE SCIENCES↗

A machine learning approach for identifying variables associated with risk of developing neutralizing antidrug antibodies to factor VIII

A key unmet need in the management of hemophilia A (HA) is the lack of clinically validated markers that are associated with the development of neutralizing antibodies to Factor VIII (FVIII) (commonly referred to as inhibitors). This study aimed to identify relevant biomarkers for FVIII inhibition using Machine Learning (ML) and Explainable AI (XAI) using the My Life Our Future (MLOF) research repository. The dataset includes biologically relevant variables such as age, race, sex, ethnicity, and the variants in the F8 gene. In addition, we previously carried out Human Leukocyte Antigen Class II (HLA-II) typing on samples obtained from the MLOF repository. Using this information, we derived other patient-specific biologically and genetically important variables. These included identifying the number of foreign FVIII derived peptides, based on the alignment of the endogenous FVIII and infused drug sequences, and the foreign-peptide HLA-II molecule binding affinity calculated using NetMHCIIpan. The data were processed and trained with multiple ML classification models to identify the top performing models. The top performing model was then chosen to apply XAI via SHAP, (SHapley Additive exPlanations) to identify the variables critical for the prediction of FVIII inhibitor development in a hemophilia A patient. Using XAI we provide a robust and ranked identification of variables that could be predictive for developing inhibitors to FVIII drugs in hemophilia A patients. These variables could be validated as biomarkers and used in making clinical decisions and during drug development. The top five variables for predicting inhibitor development based on SHAP values are: (i) the baseline activity of the FVIII protein, (ii) mean affinity of all foreign peptides for HLA DRB 3, 4, & 5 alleles, (iii) mean affinity of all foreign peptides for HLA DRB1 alleles), (iv) the minimum affinity among all foreign peptides for HLA DRB1 alleles, and (v) F8 mutation type.

60 APPLIED LIFE SCIENCES↗

Identifying human failure events (HFEs) for external hazard probabilistic risk assessment

In recent years, several advancements in nuclear power plant (NPP) probabilistic risk assessment (PRA) have been driven by increased understanding of external hazards, plant response, and uncertainties. However, major sources of uncertainty associated with external hazard PRA remain. One important source is how risk-significant human actions that are carried out to enable plant response and recovery from natural hazards cause the close coupling of physical impacts on plants and overall plant risk during these hazard events. This makes human reliability and human-plant interactions important elements to consider in resolving PRA gaps in external hazards. One of the challenges in considering human response in external hazard probabilistic risk assessment (XHPRA) is that most existing human reliability analysis (HRA) models were not developed for assessing actions outside the control room (termed ex-control room actions) and hazard response. To support this new scope, HRA models will need to be developed or modified to support identification of human activities, causal factors, and uncertainties inherent in external hazard response, thereby providing insights regarding event timing and physical event conditions as they relate to human performance. In this study, there are two main objectives: (1) evaluate the applicability of an existing cognitive-based HRA method, Phoenix, to ex-control room actions, and (2) identify sources of uncertainty to be characterized or reduced in order to make this method suitable for XHPRA. The first step of such work is performed by assessing the suitability of existing HRA methods to support identifying human failure events (HFEs) for human response to flooding hazards. These HFEs are human actions or inactions that are involved in human responses to flooding hazards and could contribute to the loss of a critical function for the plant in the scenario being examined. Here, in this work, decomposition analyses using the cognitive-based Phoenix HRA model are used to identify HFEs. The Phoenix method was found to be suitable for analyzing ex-control room actions as well as identifying specific HFEs and underlying crew failure modes (CFMs). However, the method's suitability for use in ex-control room actions would benefit from expanding the available CFMs to accommodate a larger variety of physical and communication tasks.

42 ENGINEERING↗

An Activity-Based Oxaziridine Platform for Identifying and Developing Covalent Ligands for Functional Allosteric Methionine Sites: Redox-Dependent Inhibition of Cyclin-Dependent Kinase 4

Activity-based protein profiling (ABPP) is a versatile strategy for identifying and characterizing functional protein sites and compounds for therapeutic development. However, the vast majority of ABPP methods for covalent drug discovery target highly nucleophilic amino acids such as cysteine or lysine. Here, we report a methionine-directed ABPP platform using Redox-Activated Chemical Tagging (ReACT), which leverages a biomimetic oxidative ligation strategy for selective methionine modification. Application of ReACT to oncoprotein cyclin-dependent kinase 4 (CDK4) as a representative high-value drug target identified three new ligandable methionine sites. We then synthesized a methionine-targeting covalent ligand library bearing a diverse array of heterocyclic, heteroatom, and stereochemically rich substituents. ABPP screening of this focused library identified 1oxF11 as a covalent modifier of CDK4 at an allosteric M169 site. This compound inhibited kinase activity in a dose-dependent manner on purified protein and in breast cancer cells. Further investigation of 1oxF11 found prominent cation-π and H-bonding interactions stabilizing the binding of this fragment at the M169 site. Quantitative mass-spectrometry studies validated 1oxF11 ligation of CDK4 in breast cancer cell lysates. Further biochemical analyses revealed cross-talk between M169 oxidation and T172 phosphorylation, where M169 oxidation prevented phosphorylation of the activating T172 site on CDK4 and blocked cell cycle progression. Finally, by identifying a new mechanism for allosteric methionine redox regulation on CDK4 and developing a unique modality for its therapeutic intervention, this work showcases a generalizable platform that provides a starting point for engaging in broader chemoproteomics and protein ligand discovery efforts to find and target previously undruggable methionine sites.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Radiogenic Noble Gas Nuclides to Identify and Characterize Rock Fracturing

Abstract Fracture‐released radiogenic noble gas nuclides are used to identify locations and constrain the volume of new fracture creation during subsurface detonations. Real‐time, in situ noble gases and reactive gases were monitored using a field‐deployed mass spectrometer and automated sampling system in a multilevel borehole array. Released gases were measured after two different detonations having distinct energy, pressure, and gas volume characteristics. Explosive‐derived gases (N 2 O, CO 2 ) and excess radiogenic 4 He and 40 Ar above atmospheric background are used to identify locations of gas transport and new fracture creation after each detonation. Fracture‐released radiogenic 4 He is used to constrain the volume of newly created fractures with a model of helium release from fracturing. Explosive by‐product gas was observed in multiple locations both near and distal to the shot locations for both detonations. Radiogenic 4 He and 40 Ar release from rock damage was observed in locations near the detonation after the second, more powerful detonation. Observed 4 He response is consistent with a model of diffusive release from newly created fractures. Volume of new fractures estimated from the 4 He release ranges from 1 to 5 m 2 with apertures ranging from 0.1 to 1 m. Our results provide evidence that radiogenic noble gases released during fracture creation can be identified at the field scale in real time and used to identify timing and location of fracture creation during deformation events. This technique could be useful in subsurface science and engineering problems where the location and amount of newly created rock fracturing is of interest including fault rupture, mine safety, subsurface detonation monitoring and reservoir stimulation.

58 GEOSCIENCES↗

Hot or Not? An Evaluation of Methods for Identifying Hot Moments of Nitrous Oxide Emissions From Soils

Abstract Effectively quantifying hot moments of nitrous oxide (N 2 O) emissions from agricultural soils is critical for managing this potent greenhouse gas. However, we are challenged by a lack of standard approaches for identifying hot moments, including (a) determining thresholds above which emissions are considered hot moments, and (b) considering seasonal variation in the magnitude and frequency distribution of net N 2 O fluxes. We used one year of hourly N 2 O flux measurements from 16 autochambers that varied in flux magnitude and frequency distribution in a conventionally tilled maize field in central Illinois, USA, to compare three approaches to identify hot moment thresholds: standard deviations (SD) above the mean, 1.5x the interquartile range (IQR), and isolation forest (IF) identification of anomalous values. We also compared these approaches on seasonally subdivided data (early, late, and non‐growing seasons) versus the whole year. Our analyses revealed that 1.5x IQR method best identified N 2 O hot moments. In contrast, using 2 or 4 SD both yielded hot moment threshold values too high, and IF yielded threshold values too low, leading to missed N 2 O hot moments or low net N 2 O fluxes mischaracterized as hot moments, respectively. Furthermore, seasonally subdividing the data set not only facilitated identification of smaller hot moments in the late‐ and non‐growing seasons when N 2 O hot moments were generally smaller but it also increased hot moment threshold values in the early growing season when N 2 O hot moments were larger. Consequently, of the methods evaluated here, we recommend using the 1.5x IQR method on whole year data sets to identify N 2 O hot moments.

Stuchiner, Emily R. [Institute for Sustainability,↗

An unsupervised machine learning based approach to identify efficient spin-orbit torque materials

Materials with large spin–orbit torque (SOT) hold considerable significance for many spintronic applications because of their potential for energy-efficient magnetization switching. Unfortunately, most of the existing materials exhibit an SOT efficiency factor that is much less than unity, requiring a large current for magnetization switching. The search for new materials that can exhibit an SOT efficiency much greater than unity is a topic of active research, and only a few such materials have been identified using conventional approaches. In this paper, we present a machine learning-based approach using a word embedding model that can identify new results by deciphering non-trivial correlations among various items in a specialized scientific text corpus. We show that such a model can be used to identify materials likely to exhibit high SOT and rank them according to their expected SOT strengths. The model captured the essential spintronics knowledge embedded in scientific abstracts within various materials science, physics, and engineering journals and identified 97 new materials to exhibit high SOT. Among them, 16 candidate materials are expected to exhibit an SOT efficiency greater than unity, and one of them has recently been confirmed with experiments with quantitative agreement with the model prediction.

Sayed, Shehrin↗

Identifying COVID-19 cases and extracting patient reported symptoms from Reddit using natural language processing

We used social media data from “covid19positive” subreddit, from 03/2020 to 03/2022 to identify COVID-19 cases and extract their reported symptoms automatically using natural language processing (NLP). We trained a Bidirectional Encoder Representations from Transformers classification model with chunking to identify COVID-19 cases; also, we developed a novel QuadArm model, which incorporates Question-answering, dual-corpus expansion, Adaptive rotation clustering, and mapping, to extract symptoms. Our classification model achieved a 91.2% accuracy for the early period (03/2020-05/2020) and was applied to the Delta (07/2021–09/2021) and Omicron (12/2021–03/2022) periods for case identification. We identified 310, 8794, and 12,094 COVID-positive authors in the three periods, respectively. The top five common symptoms extracted in the early period were coughing (57%), fever (55%), loss of sense of smell (41%), headache (40%), and sore throat (40%). During the Delta period, these symptoms remained as the top five symptoms with percent authors reporting symptoms reduced to half or fewer than the early period. During the Omicron period, loss of sense of smell was reported less while sore throat was reported more. Our study demonstrated that NLP can be used to identify COVID-19 cases accurately and extracted symptoms efficiently.

60 APPLIED LIFE SCIENCES↗

Pan-cancer proteogenomic investigations identify post-transcriptional kinase targets

Identifying genomic alterations of cancer proteins has guided the development of targeted therapies, but proteomic analyses are required to validate and reveal new treatment opportunities. Herein, we develop a new algorithm, OPPTI, to discover overexpressed kinase proteins across 10 cancer types using global mass spectrometry proteomics data of 1,071 cases. OPPTI outperforms existing methods by leveraging multiple co-expressed markers to identify targets overexpressed in a subset of tumors. OPPTI-identified overexpression of ERBB2 and EGFR proteins correlates with genomic amplifications, while CDK4/6, PDK1, and MET protein overexpression frequently occur without corresponding DNA- and RNA-level alterations. Analyzing CRISPR screen data, we confirm expression-driven dependencies of multiple currently-druggable and new target kinases whose expressions are validated by immunochemistry. Identified kinases are further associated with up-regulated phosphorylation levels of corresponding signaling pathways. Collectively, our results reveal protein-level aberrations—sometimes not observed by genomics—represent cancer vulnerabilities that may be targeted in precision oncology.

60 APPLIED LIFE SCIENCES↗

Identifiability and characterization of transmon qutrits through Bayesian experimental design

Robust control of a quantum system is essential to utilize the current noisy quantum hardware to its full potential, such as quantum algorithms. To achieve such a goal, a systematic search for an optimal control for any given experiment is essential. The design of optimal control pulses requires accurate numerical models and, therefore, accurate characterization of the system parameters. We present an online Bayesian approach for quantum characterization of qutrit systems, which automatically and systematically identifies optimal experiments that provide maximum information on the system parameters, thereby greatly reducing the number of experiments that need to be performed on the quantum testbed. Unlike most characterization protocols that provide point-estimates of the parameters, the proposed approach is able to estimate their probability distribution. The applicability of the Bayesian experimental design technique was demonstrated on test problems, where each experiment was defined by a parameterized control pulse. In addition to this, we also present an approach for iterative pulse extension, which is robust under uncertainties in transition frequencies and coherence times, and shot noise, despite being initialized with wide uninformative priors. Furthermore, we provide a mathematical proof of the theoretical identifiability of the model parameters and present conditions on the quantum state under which the parameters are identifiable. The proof and conditions for identifiability are presented for both closed and open quantum systems using the Schrödinger equation and the Lindblad master equation, respectively.

97 MATHEMATICS AND COMPUTING↗

AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency Analysis

Checkpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 13 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint.

HPC↗

An Information Theoretic Approach to Identify Dominant Voltage Influencers for Unbalanced Distribution Systems

Smart distribution grid with multiple renewable energy sources can experience random voltage fluctuations due to variable generation, which may result in voltage violations. Traditional voltage control algorithms are inadequate to handle fast voltage variations. Therefore, new dynamic control methods are being developed that can significantly benefit from the knowledge of dominant voltage influencer (DVI) nodes. DVI nodes for a particular node of interest refer to nodes that have a relatively high impact on the voltage fluctuations at that node. Conventional power flow-based algorithms to identify DVI nodes are computationally complex, which limits their use in real-time applications. This paper proposes a novel information theoretic voltage influencing score (VIS) that quantifies the voltage influencing capacity of nodes with DERs/active loads in a three phase unbalanced distribution system. VIS is then employed to rank the nodes and identify the DVI set. VIS is derived analytically in a computationally efficient manner and its efficacy to identify DVI nodes is validated using the IEEE 37-node test system. It is shown through experiments that KL divergence and Bhattacharyya distance are effective indicators of DVI nodes with an identifying accuracy of more than 90%. Additionally, the computation burden is also reduced by an order of 5, thus providing the foundation for efficient voltage control.

42 ENGINEERING↗

Phylogeography of the blacklegged tick ( Ixodes scapularis ) throughout the USA identifies candidate loci for differences in vectorial capacity

Abstract The blacklegged tick ( Ixodes scapularis ( Journal of the Academy of Natural Sciences of Philadelphia , 1821, 2 , 59)) is a vector of Borrelia burgdorferi sensu stricto ( s.s .) ( International Journal of Systematic Bacteriology , 1984, 34 , 496), the causative bacterial agent of Lyme disease, part of a slow‐moving epidemic of Lyme borreliosis spreading across the northern hemisphere. Well‐known geographical differences in the vectorial capacity of these ticks are associated with genetic variation. Despite the need for detailed genetic information in this disease system, previous phylogeographical studies of these ticks have been restricted to relatively few populations or few genetic loci. Here we present the most comprehensive phylogeographical study of genome‐wide markers in I. scapularis , conducted by using 3RAD (triple‐enzyme restriction‐site associated sequencing) and surveying 353 ticks from 33 counties throughout the species' range. We found limited genetic variation among populations from the Northeast and Upper Midwest, where Lyme disease is most common, and higher genetic variation among populations from the South. We identify five spatially associated genetic clusters of I. scapularis . In regions where Lyme disease is increasing in frequency, the I. scapularis populations genetically group with ticks from historically highly Lyme‐endemic regions. Finally, we identify 10 variable DNA sites that contribute the most to population differentiation. These variable sites cluster on one of the chromosome‐scale scaffolds for I. scapularis and are within identified genes. Our findings illuminate the need for additional research to identify loci causing variation in the vectorial capacity of I. scapularis and where additional tick sampling would be most valuable to further understand disease trends caused by pathogens transmitted by I. scapularis .

3RAD↗

Deep learning-enabled natural language processing to identify directional pharmacokinetic drug–drug interactions

Background. During drug development, it is essential to gather information about the change of clinical exposure of a drug (object) due to the pharmacokinetic (PK) drug-drug interactions (DDIs) with another drug (precipitant). While many natural language processing (NLP) methods for DDI have been published, most were designed to evaluate if (and what kind of) DDI relationships exist in the text, without identifying the direction of DDI (object vs. precipitant drug). Here we present a method for the automatic identification of the directionality of a PK DDI from literature or drug labels. Methods. We reannotated the Text Analysis Conference (TAC) DDI track 2019 corpus for identifying the direction of a PK DDI and evaluated the performance of a fine-tuned BioBERT model on this task by following the training and validation steps prespecified by TAC. Results. This initial attempt showed the model achieved an F-score of 0.82 in identifying sentences as containing PK DDI and an F-score of 0.97 in identifying object versus precipitant drugs in those sentences. Discussion and conclusion. Despite a growing list of NLP methods for DDI extraction, most of them use a common set of corpora to perform general purpose tasks (e.g., classifying a sentence into one of several fixed DDI categories). There is a lack of coordination between the drug development and biomedical informatics method development community to develop corpora and methods to perform specific tasks (e.g., extract clinical exposure changes due to PK DDI). We hope that our effort can encourage such a coordination so that more “fit for purpose” NLP methods could be developed and used to facilitate the drug development process.

59 BASIC BIOLOGICAL SCIENCES↗