Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AuC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Probabilistic Hydrological Estimation of LandSlides (PHELS): Global Ensemble Landslide Hazard Modelling

In this study we present a model for the global Probabilistic Hydrological Estimation of LandSlides (PHELS). PHELS estimates the daily hazard of hydrologically triggered landslides at a coarse spatial resolution of 36 km, by combining landslide susceptibility (LSS) and (percentiles of) hydrological variable(s). The latter include daily rainfall, a 7-day antecedent rainfall index (ARI7) or root-zone soil moisture content (rzmc) as hydrological predictor variables, or the combination of rainfall and rzmc. The hazard estimates with any of these predictor variables have areas under the receiver operating characteristic curve (AUC) above 0.68. The best performance was found with combined rainfall and rzmc predictors (AUC = 0.79), which resulted in the lowest number of missed alarms (especially during spring) and false alarms. Furthermore, PHELS provides hazard uncertainty estimates by generating ensemble simulations based on repeated sampling of LSS and the hydrological predictor variables. The estimated hazard uncertainty follows the behaviour of the input variable uncertainties, is about 13.6 % of the estimated hazard value on average across the globe and in time and is smallest for very low and very high hazard values.

Anne Felsberg↗

Machine Learning for Gearbox Fault Prediction by Using Both Scada and Modeled Data

This presentation outlines the work in the paper titled "Prognostics of Wind Turbine Gearbox Bearing Failures Using SCADA and Modeled Data" published by the PHM Society and presented at its 2020 annual conference. It is accessible at https://papers.phmsociety.org/index.php/phmconf/article/download/1292/862. The technical work is on machine learning approaches for prognostics for gearbox faults. The methodology combines SCADA time series data and physics domain modeling data, derived from the models developed by the NREL team, as inputs to machine learning models to predict gearbox bearing failures with one month lead time. Based on SCADA data, modeled data, and bearing failure log data from an actual wind plant, the performances of different machine-learning models on unseen data are then evaluated using industry-standard metrics such as precision, recall, and F1 score, and AUC (area under receiver operating characteristic curve). Results show the overall system performance enhancement in predicting bearing failure when modeled data are included with SCADA data. The reduction in terms of false alarms is about 50%, and improvement in terms of precision, and F1 score, and AUC is about 33%, and 12%, and 6% respectively, based on the best modeling case in this study.

49 EE - Wind and Water Power Program - Wind (EE-4W↗

Data-driven modeling of dynamic occupant thermostat override behavior for demand response applications

Buildings consume nearly 40% of global energy and produce similar emissions. Whiletechnological advances address efficiency, occupant behavior causes energy use variations up to 300% between identical buildings. This gap between predicted and actual building performance impacts building design, operations, and grid demand management programs. Through analyses of smart thermostat data from 1,400 single-occupant homes, the researchdemonstrates that occupants respond to 8°F thermostat setpoint changes within a median of 15 minutes, while 2°F changes trigger responses within a median of 30 minutes. This highlights an understudied temporal relationship between thermostat setbacks and response time of occupant behaviors. Models of such behavior dynamics are required to incorporate occupant impacts into building performance simulation. A key contribution of this dissertation is the Thermal Frustration Theory (TFT), which positsthat thermal discomfort driven behaviors are caused by the time-accumulation of discomfort, not simply a temperature deviation threshold or a delay from an initiating event. Using a dataset of 634 thermostats, each with 25+ manual setpoint changes, a comparative analysis of TFT and comfort zone and a delayed response theories demonstrated that personalized TFT models better predict when manual setpoint change occur. This was measured by the area under the curve statistical measure (AUC); all three models perform similarly by a Matthews Correlation Coefficient measure. Higher AUC performance is especially important for modeling occupant behavior in demand response programs where false negatives of rare occupant interactions could adversely affect grid stability. EnergyPlus based simulations were conducted with TFT-derived occupant models, demonstrating the ability to identify parameters of known TFT models from only data observable with smart thermostats, even under the presence of noise from routine overrides. Overall, the dissertation highlights that thermostat interactions are neither static,instantaneous, nor driven solely by the environment. Instead, temporal accumulation of discomfort and routine-based behavior play important roles. The methodology and results offer a pathway towards more accurate modeling of human-building interactions for policy assessment, building design, and demand response programs.

Sharma, Kunind [Northeastern University] (ORCID:00↗

Enhancing Automotive Intrusion Detection Through Multi-Modal Fusion: A CAN FD-LiDAR Approach

As vehicles become smarter and more autonomous, they increasingly depend on advanced sensors and communication technologies to operate securely. However, such growing dependence on technology—whether it’s CAN (Controller Area Network) for internal communication or LiDAR (Light Detection and Ranging) for sensing the world around them—also expands the attack surface for the types of cyber attacks. Traditional intrusion detection systems (IDS) typically monitor these systems in isolation, limiting their ability to detect sophisticated, crosssystem attacks. To address this, we propose a multi-modal fusion approach that combines real-world CAN FD signals (from the HCRL dataset) with LiDAR features (from the nuScenes dataset) to enhance attack detection. Our method employs a twostage ensemble approach. Calibrated XGBoost and LightGBM models initially process CAN FD (Fuzzing Data) and LiDAR data independently, detecting timing anomalies and space abnormalities. They are subsequently logarithmically combined with a logistic regression meta-model along with 17 engineered features capturing cross-modal behavior, prediction conflicts, and nonlinear interactions. This approach achieves an AUC of 0.87 and an F1-score of 0.82, surpassing single-modality baselines and early fusion methods, at merely 2 ms inference latency. Compared with deep learning competitors, it is 3 times more efficient, providing a lightweight, interpretable, and real time solution to automotive cybersecurity.

97 MATHEMATICS AND COMPUTING↗

Evaluation of Plasma Phosphorylated Tau217 for Differentiation Between Alzheimer Disease and Frontotemporal Lobar Degeneration Subtypes Among Patients With Corticobasal Syndrome

Plasma phosphorylated tau217 (p-tau217), a biomarker of Alzheimer disease (AD), is of special interest in corticobasal syndrome (CBS) because autopsy studies have revealed AD is the driving neuropathology in up to 40% of cases. This differentiates CBS from other 4-repeat tauopathy (4RT)–associated syndromes, such as progressive supranuclear palsy Richardson syndrome (PSP-RS) and nonfluent primary progressive aphasia (nfvPPA), where underlying frontotemporal lobar degeneration (FTLD) is typically the primary neuropathology. To validate plasma p-tau217 against positron emission tomography (PET) in 4RT-associated syndromes, especially CBS. This multicohort study with 6, 12, and 24-month follow-up recruited adult participants between January 2011 and September 2020 from 8 tertiary care centers in the 4RT Neuroimaging Initiative (4RTNI). All participants with CBS (n = 113), PSP-RS (n = 121), and nfvPPA (n = 39) were included; other diagnoses were excluded due to rarity (n = 29). Individuals with PET-confirmed AD (n = 54) and PET-negative cognitively normal control individuals (n = 59) were evaluated at University of California San Francisco. Operators were blinded to the cohort. Plasma p-tau217, measured by Meso Scale Discovery electrochemiluminescence, was validated against amyloid-β (Aβ) and flortaucipir (FTP) PET. Imaging analyses used voxel-based morphometry and bayesian linear mixed-effects modeling. Clinical biomarker associations were evaluated using longitudinal mixed-effect modeling. Of 386 participants, 199 (52%) were female, and the mean (SD) age was 68 (8) years. Plasma p-tau217 was elevated in patients with CBS with positive Aβ PET results (mean [SD], 0.57 [0.43] pg/mL) or FTP PET (mean [SD], 0.75 [0.30] pg/mL) to concentrations comparable to control individuals with AD (mean [SD], 0.72 [0.37]), whereas PSP-RS and nfvPPA showed no increase relative to control. Within CBS, p-tau217 had excellent diagnostic performance with area under the receiver operating characteristic curve (AUC) for Aβ PET of 0.87 (95% CI, 0.76-0.98; P < .001) and FTP PET of 0.93 (95% CI, 0.83-1.00; P < .001). At baseline, individuals with CBS-AD (n = 12), defined by a PET-validated plasma p-tau217 cutoff 0.25 pg/mL or greater, had increased temporoparietal atrophy at baseline compared to individuals with CBS-FTLD (n = 39), whereas longitudinally, individuals with CBS-FTLD had faster brainstem atrophy rates. Individuals with CBS-FTLD also progressed more rapidly on a modified version of the PSP Rating Scale than those with CBS-AD (mean [SD], 3.5 [0.5] vs 0.8 [0.8] points/year; P = .005). In this cohort study, plasma p-tau217 had excellent diagnostic performance for identifying Aβ or FTP PET positivity within CBS with likely underlying AD pathology. Plasma P-tau217 may be a useful and inexpensive biomarker to select patients for CBS clinical trials.

59 BASIC BIOLOGICAL SCIENCES↗

Automated detection of mild cognitive impairment and dementia from voice recordings: A natural language processing approach

Automated computational assessment of neuropsychological tests would enable widespread, cost-effective screening for dementia. A novel natural language processing approach is developed and validated to identify different stages of dementia based on automated transcription of digital voice recordings of subjects’ neuropsychological tests conducted by the Framingham Heart Study (n = 1084). Transcribed sentences from the test were encoded into quantitative data and several models were trained and tested using these data and the participants’ demographic characteristics. Average area under the curve (AUC) on the held-out test data reached 92.6%, 88.0%, and 74.4% for differentiating Normal cognition from Dementia, Normal or Mild Cognitive Impairment (MCI) from Dementia, and Normal from MCI, respectively. The proposed approach offers a fully automated identification of MCI and dementia based on a recorded neuropsychological test, providing an opportunity to develop a remote screening tool that could be adapted easily to any language.

neuropsychological↗

Feature‐shared adaptive‐boost deep learning for invasiveness classification of pulmonary subsolid nodules in CT images

Purpose In clinical practice, invasiveness is an important reference indicator for differentiating the malignant degree of subsolid pulmonary nodules. These nodules can be classified as atypical adenomatous hyperplasia (AAH), adenocarcinoma in situ (AIS), minimally invasive adenocarcinoma (MIA), or invasive adenocarcinoma (IAC). The automatic determination of a nodule's invasiveness based on chest CT scans can guide treatment planning. However, it is challenging, owing to the insufficiency of training data and their interclass similarity and intraclass variation. To address these challenges, we propose a two‐stage deep learning strategy for this task: prior‐feature learning followed by adaptive‐boost deep learning. Methods The adaptive‐boost deep learning is proposed to train a strong classifier for invasiveness classification of subsolid nodules in chest CT images, using multiple 3D convolutional neural network (CNN)‐based weak classifiers. Because ensembles of multiple deep 3D CNN models have a huge number of parameters and require large computing resources along with more training and testing time, the prior‐feature learning is proposed to reduce the computations by sharing the CNN layers between all weak classifiers. Using this strategy, all weak classifiers can be integrated into a single network. Results Tenfold cross validation of binary classification was conducted on a total of 1357 nodules, including 765 noninvasive (AAH and AIS) and 592 invasive nodules (MIA and IAC). Ablation experimental results indicated that the proposed binary classifier achieved an accuracy of with an AUC of 81.3 . These results are superior compared to those achieved by three experienced chest imaging specialists who achieved an accuracy of , , and , respectively. About 200 additional nodules were also collected. These nodules covered 50 cases for each category (AAH, AIS, MIA, and IAC, respectively). Both binary and multiple classifications were performed on these data and the results demonstrated that the proposed method definitely achieves better performance than the performance achieved by nonensemble deep learning methods. Conclusions It can be concluded that the proposed adaptive‐boost deep learning can significantly improve the performance of invasiveness classification of pulmonary subsolid nodules in CT images, while the prior‐feature learning can significantly reduce the total size of deep models. The promising results on clinical data show that the trained models can be used as an effective lung cancer screening tool in hospitals. Moreover, the proposed strategy can be easily extended to other similar classification tasks in 3D medical images.

Wang, Jun↗

Trigger Detection for the sPHENIX Experiment via Bipartite Graph Networks with Set Transformer

Trigger (interesting events) detection is crucial to high-energy and nuclear physics experiments because it improves data acquisition efficiency. It also plays a vital role in facilitating the downstream offline data analysis process. The sPHENIX detector, located at the Relativistic Heavy Ion Collider in Brookhaven National Laboratory, is one of the largest nuclear physics experiments on a world scale and is optimized to detect physics processes involving charm and beauty quarks. Furthermore, these particles are produced in collisions involving two proton beams, two gold nuclei beams, or a combination of the two and give critical insights into the formation of the early universe. This paper presents a model architecture for trigger detection with geometric information from two fast silicon detectors. Transverse momentum is introduced as an intermediate feature from physics heuristics. We also prove its importance through our training experiments. Each event consists of tracks and can be viewed as a graph. A bipartite graph neural network is integrated with the attention mechanism to design a binary classification model. Compared with the state-of-the-art algorithm for trigger detection, our model is parsimonious and increases the accuracy and the AUC score by more than 15%.

97 MATHEMATICS AND COMPUTING↗

A data-driven framework for predicting machining stability: employing simulated data, operational modal analysis, and enhanced transfer learning

Chatter, a self-excited vibration phenomenon, presents a significant challenge in machining operations, particularly in high-speed milling, where it can degrade tool life, reduce material removal efficiency, and compromise workpiece quality. Addressing this challenge requires a reliable predictive model that can accommodate the complex dynamics of various machining scenarios. This study introduces a novel, data-driven approach to predicting machining stability, leveraging over 140,000 simulated datasets and employing advanced techniques such as operational modal analysis (OMA), enhanced transfer learning (TL), and receptance coupling substructure analysis (RCSA). By integrating these methodologies, the framework effectively classifies and predicts chatter across diverse operational modes, achieving robust and accurate outcomes. Our model utilizes a Random Forest (RF) classifier trained with the comprehensive dataset, which demonstrates substantial improvements in both predictive accuracy and robustness. Specifically, the RF model achieved an accuracy rate of 85%, an area under the curve (AUC) of 0.90, and an F1 score of 0.88, underscoring its capability to adapt to varying machining configurations. These results highlight the framework’s potential to enhance operational efficiency and machining quality by providing reliable chatter predictions across a broad range of machining parameters. In conclusion, this research thus offers a significant advancement in predictive maintenance for machining processes, enabling more stable and efficient manufacturing operations.

42 ENGINEERING↗

Domain Shift Analysis in Chest Radiographs Classification in a Veterans Healthcare Administration Population

This study aims to assess the impact of domain shift on chest X-ray classification accuracy and to analyze the influence of ground truth label quality and demographic factors such as age group, sex, and study year. We used a DenseNet121 model pre-trained MIMIC-CXR dataset for deep learning-based multi-label classification using ground truth labels from radiology reports extracted using the CheXpert and CheXbert Labeler. We compared the performance of the 14 chest X-ray labels on the MIMIC-CXR and Veterans Healthcare Administration chest X-ray dataset (VA-CXR). The validation of ground truth and the assessment of multi-label classification performance across various NLP extraction tools revealed that the VA-CXR dataset exhibited lower disagreement rates than the MIMIC-CXR datasets. Additionally, there were notable differences in AUC scores between models utilizing CheXpert and CheXbert. When evaluating multi-label classification performance across different datasets, minimal domain shift was observed in the unseen VA dataset, except for the label “Enlarged Cardiomediastinum.” The subgroup with the most significant variations in multi-label classification performance was study year. These findings underscore the importance of considering domain shift in chest X-ray classification tasks, paying particular attention to the temporality of the exam. Our study reveals the significant impact of domain shift and demographic factors on chest X-ray classification, emphasizing the need for improved transfer learning and robust model development. Addressing these challenges is crucial for advancing medical imaging research and improving patient care.

chest X-ray image classification↗

In-situ sensor monitoring of multi-class gas porosity formation in laser powder bed fusion using convolutional neural network

In-situ monitoring of defect formation remains a significant challenge in the laser powder bed fusion (LPBF) process. Recent advances have enabled real-time defect detection with machine learning and in-situ sensing technologies; however, most studies focus on binary classification of keyhole pores, limiting nuanced multi-class pore differentiation and formation mechanisms. This work introduces a multi-class pore detection framework (no pore, small pores < 15 µm, and large pores > 15 µm) by leveraging photodiode sensor data alongside high-fidelity synchrotron X-ray imaging. The 15 µm threshold is selected to distinguish between two fundamentally different defect mechanisms, following the physical size-mechanism boundary established by prior high-resolution synchrotron X-ray characterization of Al6061 LPBF. Distinguishing these classes is critical because large keyhole pores are structurally detrimental, whereas small gas pores are often benign, requiring different process control strategies. Thermal emission monitoring data collected simultaneously with high-speed X-ray imaging at the Stanford Synchrotron Radiation Lightsource (SSRL), are correlated with subsurface melt pool dynamics to establish ground truth. Continuous Wavelet Transform (CWT) with optimized parameters converts the photodiode time-series signals into time–frequency images, facilitating feature extraction. Convolutional Neural Networks (CNN) are then applied for real-time multi-class pore classification in an average inference time of 1 ms per signal window. It achieves 79% accuracy and an Area Under the Receiver Operating Characteristic curve (AUC ROC) score of 0.89 with five-fold cross-validation. The results demonstrate that coupling CWT-based feature engineering with CNN architecture enables reliable multi-class pore detection in Al6061 builds using affordable in-situ sensors. This approach advances scalable and affordable quality assurance in additive manufacturing by moving beyond binary defect detection toward more nuanced classification of porosity mechanisms with in-situ sensors and machine learning.

Laser powder bed fusion, Multi-class pores, In-sit↗

Analysis and prediction of intersection traffic violations using automated enforcement system data

We report that the automated enforcement system (AES) is an effective way of supplementing traditional traffic enforcement, and the traffic violation data from AES can also be effectively used for safety research. In this study, traffic violation data were used to analyze the influencing factors associated with traffic violations and to predict the probability of violations at intersections. The potential factors influencing violations include 24 independent factors related to time, space, traffic and weather. Results from a logistic model showed that the midday period, weekends, residential districts, collector roads, congested traffic conditions, high traffic flow, lower wind speed and low temperature would increase the probability of traffic violations. The probability of violations was predicted by the random forest algorithm, which was proven to be the best traffic violation prediction model among logistic regression, Gaussian naive Bayes, and support vector machine. Moreover, the proximity weighted synthetic oversampling technique (ProWSyn) method was applied to reduce the impact of the imbalance ratio (IR) and improve the model’s prediction performance. The receiver operating characteristics (ROC) curves and Precision-Recall (PR) curves illustrated that the random forest algorithm using oversampling data had the best classifier prediction performance than undersampling data. The area under curve (AUC) and out-of-bag (OOB) error with IR = 1 reached 0.914 and 0.0787, which showed the better performance of the random forest algorithm using ProWSyn in dealing with imbalanced traffic violation data.

42 ENGINEERING↗

Executive Summary of the American Radium Society Appropriate Use Criteria for Operable Esophageal and Gastroesophageal Junction Adenocarcinoma: Systematic Review and Guidelines

Limited guidance exists regarding the relative effectiveness of treatment options for nonmetastatic, operable patients with adenocarcinoma of the esophagus or gastroesophageal junction (GEJ). In this systematic review, the American Radium Society (ARS) gastrointestinal expert panel convened to develop Appropriate Use Criteria (AUC) evaluating how neoadjuvant and/or adjuvant treatment regimens compared with each other, surgery alone, or definitive chemoradiation in terms of response to therapy, quality of life, and oncologic outcomes.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonization

Objective: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. Methods: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. Results: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. Conclusions: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes.

Zhou, Doudou↗

Deep learning for Alzheimer's disease: Mapping large-scale histological tau protein for neuroimaging biomarker validation

Abnormal tau inclusions are hallmarks of Alzheimer's disease and predictors of clinical decline. Several tau PET tracers are available for neurodegenerative disease research, opening avenues for molecular diagnosis in vivo. However, few have been approved for clinical use. Understanding the neurobiological basis of PET signal validation remains problematic because it requires a large-scale, voxel-to-voxel correlation between PET and (immuno) histological signals. Large dimensionality of whole human brains, tissue deformation impacting co-registration, and computing requirements to process terabytes of information preclude proper validation. We developed a computational pipeline to identify and segment particles of interest in billion-pixel digital pathology images to generate quantitative, 3D density maps. The proposed convolutional neural network for immunohistochemistry samples, IHCNet, is at the pipeline's core. We have successfully processed and immunostained over 500 slides from two whole human brains with three phospho-tau antibodies (AT100, AT8, and MC1), spanning several terabytes of images. Our artificial neural network estimated tau inclusion from brain images, which performs with ROC AUC of 0.87, 0.85, and 0.91 for AT100, AT8, and MC1, respectively. Introspection studies further assessed the ability of our trained model to learn tau-related features. We present an end-to-end pipeline to create terabytes-large 3D tau inclusion density maps co-registered to MRI as a means to facilitate validation of PET tracers.

60 APPLIED LIFE SCIENCES↗

Identifying schools at high-risk for elevated lead in drinking water using only publicly available data

Estimating the risk of lead contamination of schools' drinking water at the State level is a complex, important, and unexplored challenge. Variable water quality among water systems and changes in water chemistry during distribution affect lead dissolution rates from pipes and fittings. In addition, the locations of lead-bearing plumbing materials are uncertain. We tested the capability of six machine learning models to predict the likelihood of lead contamination of drinking water at the schools' taps using only publicly available datasets. The predictive features used in the models correspond to those with a proven correlation to the dominant, but commonly unavailable, factors that govern lead leaching: the presence of lead-bearing plumbing materials and water quality conducive to lead corrosion. By combining water chemistry data from public reports, socioeconomic information from the US census, and spatial features using Geographic Information Systems, we trained and tested models to estimate the likelihood of lead contaminated tap water in over 8,000 schools across California and Massachusetts. Our best-performing model was a Random Forest, with a 10-fold cross validation score of 0.88 for Massachusetts and 0.78 for California using the average Area Under the Receiver Operating Characteristic Curve (ROC AUC) metric. The model was then used to assign a lead leaching risk category to half of the schools across California (the other half was used for training). There was good agreement between the modeled risk categories and the actual lead leaching outcomes for every school; however, the model overestimated the lead leaching risk in up to 17% of the schools. This model is the first of its kind to offer a tool to predict the risk of lead leaching in schools at the State level. Further use of this model can help deploy limited resources more effectively to prevent childhood lead exposure from school drinking water.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Benzo[$\mathcal{a}$]pyrene toxicokinetics in humans following dietary supplementation with 3,3'-diindolylmethane (DIM) or Brussels sprouts

Utilizing the atto-zeptomole sensitivity of UPLC-accelerator mass spectrometry (UPLC-AMS), we previously demonstrated significant first-pass metabolism following escalating (25–250 ng) oral micro-dosing in humans of [ 14 C]-benzo[a]pyrene ([ 14 C]-BaP). The present study examines the potential for supplementation with Brussels sprouts (BS) or 3,3' -diindolylmethane (DIM) to alter plasma levels of [ 14 C]-BaP and metabolites over a 48-h period following micro-dosing with 50 ng (5.4 nCi) [ 14 C]-BaP. Volunteers were dosed with [ 14 C]-BaP following fourteen days on a cruciferous vegetable restricted diet, or the same diet supplemented for seven days with 50 g of BS or 300 mg of BR-DIM ® prior to dosing. BS or DIM reduced total [ 14 C] recovered from plasma by 56–67% relative to nonintervention. Dietary supplementation with DIM markedly increased T max and reduced C max for [ 14 C]-BaP indicative of slower absorption. Further, both dietary treatments significantly reduced C max values of four downstream BaP metabolites, consistent with delaying BaP absorption. Dietary treatments also appeared to reduce the T 1/2 and the plasma AUC( 0,$\infty$ ) for Unknown Metabolite C, indicating some effect in accelerating clearance of this metabolite. Toxicokinetic constants for other metabolites followed the pattern for [ 14 C]-BaP (metabolite profiles remained relatively consistent) and non-compartmental analysis did not indicate other significant alterations. Significant amounts of metabolites in plasma were at the bay region of [ 14 C]-BaP irrespective of treatment. Although the number of subjects and large interindividual variation are limitations of this study, it represents the first human trial showing dietary intervention altering toxicokinetics of a defined dose of a known human carcinogen.

3,3'-diindolylmethane↗

Nanomaterial Synthesis Insights from Machine Learning of Scientific Articles by Extracting, Structuring, and Visualizing Knowledge

Nanomaterials of varying compositions and morphologies are of interest for many applications from catalysis to optics, but the synthesis of nanomaterials and their scale-up are most often time-consuming and Edisonian processes. Information gleaned from the scientific literature can help inform and accelerate nanomaterials development, but again, searching the literature and digesting the information are time-consuming manual processes for researchers. To help address these challenges, here we developed scientific article-processing tools that extract and structure information from the text and figures of nanomaterials articles, thereby enabling the creation of a personalized knowledgebase for nanomaterials synthesis that can be mined to help inform further nanomaterials development. Starting with a corpus of ~35k nanomaterials-related articles, we developed models to classify articles according to the nanomaterial composition and morphology, extract synthesis protocols from within the articles’ text, and extract, normalize, and categorize chemical terms within synthesis protocols. We demonstrate the efficiency of the proposed pipeline on an expert-labeled set of nanomaterials synthesis articles, achieving 100% accuracy on composition prediction, 95% accuracy on morphology prediction, 0.99 AUC on protocol identification, and up to a 0.87 F1-score on chemical entity recognition. In addition to processing articles’ text, microscopy images of nanomaterials within the articles are also automatically identified and analyzed to determine the nanomaterials’ morphologies and size distributions. To enable users to easily explore the database, we developed a complementary browser-based visualization tool that provides flexibility in comparing across subsets of articles of interest. We use these tools and information to identify trends in nanomaterials synthesis, such as the correlation of certain reagents with various nanomaterial morphologies, which is useful in guiding hypotheses and reducing the potential parameter space during experimental design.

36 MATERIALS SCIENCE↗