Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “AuC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

High-frequency Electrocardiogram Analysis in the Ability to Predict Reversible Perfusion Defects during Adenosine Myocardial Perfusion Imaging

Background: A previous study has shown that analysis of high-frequency QRS components (HF-QRS) is highly sensitive and reasonably specific for detecting reversible perfusion defects on myocardial perfusion imaging (MPI) scans during adenosine. The purpose of the present study was to try to reproduce those findings. Methods: 12-lead high-resolution electrocardiogram recordings were obtained from 100 patients before (baseline) and during adenosine Tc-99m-tetrofosmin MPI tests. HF-QRS were analyzed regarding morphology and changes in root mean square (RMS) voltages from before the adenosine infusion to peak infusion. Results: The best area under the curve (AUC) was found in supine patients (AUC=0.736) in a combination of morphology and RMS changes. None of the measurements, however, were statistically better than tossing a coin (AUC=0.5). Conclusion: Analysis of HF-QRS was not significantly better than tossing a coin for determining reversible perfusion defects on MPI scans.

Tragardh, Elin↗

Effect of Pharmacologically-Induced Hypovolemia on Aerobic Capacity

Decreased peak oxygen consumption (VO2pk) and an elevated exercise heart rate (HR) response are associated with a reduction in plasma volume (PV) after space flight and bed rest, a space flight analog. Reduced VO2pk and submaximal exercise tolerance would negatively impact an astronaut s ability to perform near maximal work that would be required in the event of an emergency. We previously have administered IV furosemide followed by a low salt diet to model PV loss and orthostatic intolerance observed after spaceflight. Purpose: To determine whether a pharmacologically-induced reduction in PV results in decreased VO2pk and elevated exercise HR response. Methods: Six subjects (5M, 1F) performed two graded peak cycle tests (work rate increased by 35 or 50 W every 3 min), once while normovolemic and once while hypovolemic. HR and expired respiratory gases were continuously measured. To induce hypovolemia, subjects were administered a single dose of IV furosemide (0.5 mg.kg-1) 30 hr before exercise testing and then consumed a low-salt diet (10 mEq.d(sup -1)). PV was measured using carbon monoxide rebreathing. Exercise HR and VO2 responses were quantified as the area under the curve (AUC) calculated over each quartile of the peak test, based on test time in the hypovolemia condition. Paired t-tests were used to test for differences in PV, VO2pk, and peak HR between conditions. Repeated-measures ANOVAs were used to test for differences in AUC between conditions. Results: PV (3.32+/-0.12 vs. 2.77+/-0.16 L, p<0.05) and VO2pk (3.30+/-0.67 vs. 2.90+/-0.57 L.min(sup -1), p<0.05) were lower during hypovolemia than during normovolemia, but peak HR was not different (187+/-5 vs. 187+/-5 bpm). The AUC for VO2 and HR was different (p<0.05) between conditions only in the highest quartile: HR was 4% higher and VO2 was 5% lower during the hypovolemia condition. Conclusion: The mean difference in VO2pk (-12%) between normovolemia and hypovolemia was similar to the mean difference in PV (-17%). Similar decreases in PV and VO2pk have been observed following short duration space flight, suggesting that pharmacologically-induced PV loss can be used to model microgravity-induced reductions in VO2pk.

Everett, Meghan E.↗

Classifying Agnostic Biosignatures using Raman, VNIR, and Elemental Data

How can we use our current wealth of terrestrial data, encompassing biogenic and abiogenic systems, to determine the distinguishing properties of life? SCOBI (Statistical Classification of Biosignature Information) uses machine learning techniques to algorithmically identify combinations of measurements that are “indicative of life”. A set of ~1000 observations, comprising elemental abundance, isotopic fractionation, VNIR reflectance, and (in progress) Raman spectra, have been assembled from existing literature and databases. The observations cover systems classified as “indicative alive” (e.g., cells, vegetation), “indicative non-alive” (e.g., fossils, teeth), “mixed indicative” (e.g., soil, pond water), or “non-indicative” (e.g., rocks, meteorites). VNIR data was preprocessed by linear interpolation from 400-2100 nm and smoothed with a Savitzky-Golay filter. To limit the amount of Earth-biochemistry-specific (non-agnostic) information included, the first five spectral features extracted were number of peaks, number of troughs, mean reflectance, mean peak width, and broadest peak width. To help further emphasize agnostic biosignatures, Earth-specific features such as chlorophylls have been manually flagged so that feature importance with and without them can be compared. Classifiers including k-nearest neighbors (KNN), Gaussian Naïve Bayes (GNB), logistic regression (LR), random forest (RF), and support vector machine (SVM) were implemented, as was a combination voting classifier. Performance metrics included false positive rates, false negative rates, and AUC with 50-50 test/train splits (Monte Carlo simulations). Key takeaways from this stage, prior to the inclusion of Raman spectra, are (1) the overall success rate of 0.933 AUC was most heavily influenced by the elemental abundance data; and (2) VNIR reflectance had the lowest classification performance with 0.52 AUC (58% of objects correctly classified). The next steps are to complete integration of Raman spectral data and to improve the approach to pre-processing and feature extraction for both types of spectral data, such as automated baseline removal, whole spectrum matching, and dimensionality reduction.

Biosignatures↗

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Impact of simulated reduced injected dose on the assessment of amyloid PET scans

To investigate the impact of reduced injected doses on the quantitative and qualitative assessment of the amyloid PET tracers [ 18 F]flutemetamol and [ 18 F]florbetaben. Cognitively impaired and unimpaired individuals (N = 250, 36% Aβ-positive) were included and injected with [ 18 F]flutemetamol (N = 175) or [ 18 F]florbetaben (N = 75). PET scans were acquired in list-mode (90–110 min post-injection) and reduced-dose images were simulated to generate images of 75, 50, 25, 12.5 and 5% of the original injected dose. Images were reconstructed using vendor-provided reconstruction tools and visually assessed for Aβ-pathology. SUVRs were calculated for a global cortical and three smaller regions using a cerebellar cortex reference tissue, and Centiloid was computed. Absolute and percentage differences in SUVR and CL were calculated between dose levels, and the ability to discriminate between Aβ- and Aβ + scans was evaluated using ROC analyses. Finally, intra-reader agreement between the reduced dose and 100% images was evaluated. At 5% injected dose, change in SUVR was 3.72% and 3.12%, with absolute change in Centiloid 3.35CL and 4.62CL, for [ 18 F]flutemetamol and [ 18 F]florbetaben, respectively. At 12.5% injected dose, percentage change in SUVR and absolute change in Centiloid were < 1.5%. AUCs for discriminating Aβ- from Aβ + scans were high (AUC ≥ 0.94) across dose levels, and visual assessment showed intra-reader agreement of > 80% for both tracers. This proof-of-concept study showed that for both [ 18 F]flutemetamol and [ 18 F]florbetaben, adequate quantitative and qualitative assessments can be obtained at 12.5% of the original injected dose. However, decisions to reduce the injected dose should be made considering the specific clinical or research circumstances.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Emergency department overcrowding and its associated factors at HARME medical emergency center in Eastern Ethiopia

Introduction: Emergency department (ED) overcrowding has become a significant concern as it can lead to compromised patient care in emergency settings. Various tools have been used to evaluate overcrowding in ED. However, there is a lack of data regarding this issue in resource-limited countries, including Ethiopia. This study aimed to validate NEDOCS, assess level of ED overcrowding and identify associated factors at HARME Medical Emergency Center, located in Hiwot Fana Comprehensive Specialized Hospital, Harar, Ethiopia. Methods: A cross-sectional study was conducted at the HARME Medical Emergency Center, Hiwot Fana Comprehensive Specialized Hospital, involving a total of 899 patients during 120 sampling intervals. The area under the receiver operating characteristic curves (AUC) was calculated to evaluate the agreement between objective and subjective assessments of ED overcrowding. A multivariable logistic regression analysis was employed to identify factors associated with ED overcrowding and statistically significant association was declared using 95% confidence level and a p-value < 0.05. Results: The interrater agreement showed a strong correlation with a Cohen's kappa (κ) of 0.80. The National Emergency Department Overcrowding Study Score demonstrated a strong association with subjective assessments from residents and case team nurses, with an AUC of 0.81 and 0.79, respectively. According to residents' perceptions, ED were considered overcrowded 65.8% of the time. Factors significantly associated with ED overcrowding included waiting time for triage (AOR: 2.24; 95% CI: 1.54–3.27), working time (AOR: 2.23; 95% CI: 1.52–3.26), length of stay (AOR: 2.40; 95% CI: 1.27–4.54), saturation level (AOR: 2.35; 95% CI: 1.31–4.20), chronic illness (AOR: 2.19; 95% CI: 1.37–3.53), and abnormal pulse rate (AOR: 1.52; 95% CI: 1.06–2.16). Conclusion: The study revealed that ED were overcrowded approximately two-thirds of the time.

60 APPLIED LIFE SCIENCES↗

Neutron scattering maps the higher-order assembly of NADPH-dependent assimilatory sulfite reductase

Precursor molecules for biomass incorporation must be imported into cells and made available to the molecular machines that build the cell. Sulfur-containing macromolecules require that sulfur be in its S2- oxidation state before assimilation into amino acids, cofactors, and vitamins that are essential to organisms throughout the biosphere. In α-proteobacteria, NADPH-dependent assimilatory sulfite reductase (SiR) performs the final six-electron reduction of sulfur. SiR is a dodecameric oxidoreductase composed of an octameric flavoprotein reductase (SiRFP) and four hemoprotein metalloenzyme oxidases (SiRHPs). SiR performs the electron transfer reduction reaction to produce sulfide from sulfite through coordinated domain movements and subunit interactions without release of partially reduced intermediates. Efforts to understand the electron transfer mechanism responsible for SiR’s efficiency are confounded by structural heterogeneity arising from intrinsically disordered regions throughout its complex, including the flexible linker joining SiRFP’s flavin-binding domains. As a result, high-resolution structures of SiR dodecamer and its subcomplexes are unknown, leaving a gap in the fundamental understanding of how SiR performs this uniquely large-volume electron transfer reaction. In this work, we use deuterium labeling, in vitro reconstitution, analytical ultracentrifugation (AUC), small-angle neutron scattering (SANS), and neutron contrast variation (NCV) to observe the relative subunit positions within SiR’s higher-order assembly. AUC and SANS reveal SiR to be a flexible dodecamer and confirm the mismatched SiRFP and SiRHP subunit stoichiometry. NCV shows that the complex is asymmetric, with SiRHP on the periphery of the complex and the centers of mass between SiRFP and SiRHP components over 100 Å apart. SiRFP undergoes compaction upon assembly into SiR’s dodecamer and SiRHP adopts multiple positions in the complex. The resulting map of SiR’s higher-order structure supports a cis/trans mechanism for electron transfer between domains of reductase subunits as well as between tightly bound or transiently interacting reductase and oxidase subunits.

59 BASIC BIOLOGICAL SCIENCES↗

Benzo[a]pyrene (BaP) metabolites predominant in human plasma following escalating oral micro-dosing with [ 14 C]-BaP

Benzo[a]pyrene (BaP) is formed by incomplete combustion of organic materials (petroleum, coal, tobacco, etc.). BaP is designated by the International Agency for Research on Cancer as a group 1 known human carcinogen; a classification supported by numerous studies in preclinical models and epidemiology studies of exposed populations. Risk assessment relies on toxicokinetic and cancer studies in rodents at doses 5–6 orders of magnitude greater than average human uptake. Using a dose–response design at environmentally relevant concentrations, this study follows uptake, metabolism, and elimination of [ 14 C]-BaP in human plasma by employing UPLC - accelerator mass spectrometry (UPLC-AMS). Volunteers were administered 25, 50, 100, and 250 ng (2.7–27 nCi) of [ 14 C]-BaP (with interceding minimum 3-week washout periods) with quantification of parent [ 14 C]-BaP and metabolites in plasma measured over 48 h. [ 14 C]-BaP median T max was 30 min with C max and area under the curve (AUC) approximating dose-dependency. Marked inter-individual variability in plasma pharmacokinetics following a 250 ng dose was seen with 7 volunteers as measured by the C max (8.99 ± 7.08 ng × mL -1 ) and AUC0-48hr (68.6 ± 64.0 fg × hr -1 × mL -1 ). Approximately 3–6% of the [ 14 C] recovered (AUC 0-48 hr ) was parent compound, demonstrating extensive metabolism following oral dosing. Metabolite profiles showed that, even at the earliest time-point (30 min), a substantial percentage of [ 14 C] in plasma was polar BaP metabolites. The best fit modeling approach identified non-compartmental apparent volume of distribution of BaP as significantly increasing as a function of dose (p = 0.004). Bay region tetrols and dihydrodiols predominated, suggesting not only was there extensive first pass metabolism but also potentially bioactivation. AMS enables the study of environmental carcinogens in humans with de minimus risk, allowing for important testing and validation of physiologically based pharmacokinetic models derived from animal data, risk assessment, and the interpretation of data from high-risk occupationally exposed populations.

59 BASIC BIOLOGICAL SCIENCES↗

Towards deep computer vision for in-line defect detection in polymer electrolyte membrane fuel cell materials

Polymer Electrolyte Membrane (PEM) fuel cells are a promising source of alternative energy. However, their production is limited by a lack of well-established methods for quality control of their constituent materials like the membrane-electrode assembly during roll-to-roll manufacturing. One potential solution is the implementation of deep learning methods to detect unwanted defects through their detection in scanned images. Here we explore the detection of defects like scratches, pinholes, and scuffs in a sample dataset of PEM optical images using two deep learning algorithms: Patch Distribution Modeling (PaDiM) for unsupervised anomaly detection and Faster-RCNN for supervised object detection. Both methods achieve scores on performance metrics (ROC-AUC and PRO-AUC for PaDiM and AP for Faster-RCNN) that are comparable to their scores on benchmark datasets. These methods also have the potential to detect a wider range of defects compared to IR thermography and previous optical inspection methods. Overall, deep learning shows promise at detecting relevant defects of interest and has the potential to achieve real-time defect detection.

30 DIRECT ENERGY CONVERSION↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Toward Guided Mutagenesis: Gaussian Process Regression Predicts MHC Class II Antigen Mutant Binding

Antigen-specific immunotherapies (ASI) require successful loading and presentation of antigen peptides into the major histocompatibility complex (MHC) binding cleft. One route of ASI design is to mutate native antigens for either stronger or weaker binding interaction to MHC. Exploring all possible mutations is costly both experimentally and computationally. To reduce experimental and computational expense, here we investigate the minimal amount of prior data required to accurately predict the relative binding affinity of point mutations for peptide-MHC class II (pMHCII) binding. Using data from different residue subsets, we interpolate pMHCII mutant binding affinities by Gaussian process (GP) regression of residue volume and hydrophobicity. We apply GP regression to an experimental data set from the Immune Epitope Database, and theoretical data sets from NetMHCIIpan and Free Energy Perturbation calculations. We find that GP regression can predict binding affinities of nine neutral residues from a six-residue subset with an average R 2 coefficient of determination value of 0.62 ± 0.04 (±95% CI), average error of 0.09 ± 0.01 kcal/mol (±95% CI), and with an receiver operating characteristic (ROC) AUC value of 0.92 for binary classification of enhanced or diminished binding affinity. Similarly, metrics increase to an R2 value of 0.69 ± 0.04, average error of 0.07 ± 0.01 kcal/mol, and an ROC AUC value of 0.94 for predicting seven neutral residues from an eight-residue subset. Our work finds that prediction is most accurate for neutral residues at anchor residue sites without register shift. This work holds relevance to predicting pMHCII binding and accelerating ASI design.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electrocardiographic changes predate Parkinson’s disease onset

Autonomic nervous system involvement precedes the motor features of Parkinson’s disease (PD). Our goal was to develop a proof-of-concept model for identifying subjects at high risk of developing PD by analysis of cardiac electrical activity. We used standard 10-s electrocardiogram (ECG) recordings of 60 subjects from the Honolulu Asia Aging Study including 10 with prevalent PD, 25 with prodromal PD, and 25 controls who never developed PD. Various methods were implemented to extract features from ECGs including simple heart rate variability (HRV) metrics, commonly used signal processing methods, and a Probabilistic Symbolic Pattern Recognition (PSPR) method. Extracted features were analyzed via stepwise logistic regression to distinguish between prodromal cases and controls. Stepwise logistic regression selected four features from PSPR as predictors of PD. The final regression model built on the entire dataset provided an area under receiver operating characteristics curve (AUC) with 95% confidence interval of 0.90 [0.80, 0.99]. The five-fold cross-validation process produced an average AUC of 0.835 [0.831, 0.839]. We conclude that cardiac electrical activity provides important information about the likelihood of future PD not captured by classical HRV metrics. Machine learning applied to ECGs may help identify subjects at high risk of having prodromal PD.

59 BASIC BIOLOGICAL SCIENCES↗

Long-term survival and second malignant tumor prediction in pediatric, adolescent, and young adult cancer survivors using Random Survival Forests: a SEER analysis

Abstract Survival and second malignancy prediction models can aid clinical decision making. Most commonly, survival analysis studies are performed using traditional proportional hazards models, which require strong assumptions and can lead to biased estimates if violated. Therefore, this study aims to implement an alternative, machine learning (ML) model for survival analysis: Random Survival Forest (RSF). In this study, RSFs were built using the U.S. Surveillance Epidemiology and End Results to (1) predict 30-year survival in pediatric, adolescent, and young adult cancer survivors; and (2) predict risk and site of a second tumor within 30 years of the first tumor diagnosis in these age groups. The final RSF model for pediatric, adolescent, and young adult survival has an average Concordance index (C-index) of 92.9%, 94.2%, and 94.4% and average time-dependent area under the receiver operating characteristic curve (AUC) at 30-years since first diagnosis of 90.8%, 93.6%, 96.1% respectively. The final RSF model for pediatric, adolescent, and young adult second malignancy has an average C-index of 86.8%, 85.2%, and 88.6% and average time-dependent AUC at 30-years since first diagnosis of 76.5%, 88.1%, and 99.0% respectively. This study suggests the robustness and potential clinical value of ML models to alleviate physician burden by quickly identifying highest risk individuals.

60 APPLIED LIFE SCIENCES↗

End-to-End Jet Classification of Boosted Top Quarks with CMS Open Data

We describe a novel application of the end-to-end deep learning technique to the task of discriminating top quark-initiated jets from those originating from the hadronization of a light quark or a gluon. The end-to-end deep learning technique combines deep learning algorithms and low-level detector representation of the high-energy collision event. In this study, we use lowlevel detector information from the simulated CMS Open Data samples to construct the top jet classifiers. To optimize classifier performance we progressively add low-level information from the CMS tracking detector, including pixel detector reconstructed hits and impact parameters, and demonstrate the value of additional tracking information even when no new spatial structures are added. Relying only on calorimeter energy deposits and reconstructed pixel detector hits, the end-to-end classifier achieves a ROC-AUC score of 0.975±0.002 for the task of classifying boosted top quark jets. After adding derived track quantities, the classifier ROC-AUC score increases to 0.9824±0.0013, serving as the first performance benchmark for these CMS Open Data samples.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Optimizing a magnitude-limited spectroscopic training sample for photometric classification of supernovae

ABSTRACT In preparation for photometric classification of transients from the Legacy Survey of Space and Time (LSST) we run tests with different training data sets. Using estimates of the depth to which the 4-m Multi-Object Spectroscopic Telescope (4MOST) Time Domain Extragalactic Survey (TiDES) can classify transients, we simulate a magnitude-limited sample reaching rAB ≈ 22.5 mag. We run our simulations with the software snmachine, a photometric classification pipeline using machine learning. The machine-learning algorithms struggle to classify supernovae when the training sample is magnitude limited, in contrast to representative training samples. Classification performance noticeably improves when we combine the magnitude-limited training sample with a simulated realistic sample of faint high-redshift supernovae observed from larger spectroscopic facilities; the algorithms’ range of average area under receiver operator characteristic curve (AUC) scores over 10 runs increases from 0.547–0.628 to 0.946–0.969 and purity of the classified sample reaches 95 per cent in all runs for two of the four algorithms. By creating new, artificial light curves using the augmentation software avocado, we achieve a purity in our classified sample of 95 per cent in all 10 runs performed for all machine-learning algorithms considered. We also reach a highest average AUC score of 0.986 with the artificial neural network algorithm. Having ‘true’ faint supernovae to complement our magnitude-limited sample is a crucial requirement in optimization of a 4MOST spectroscopic sample. However, our results are a proof of concept that augmentation is also necessary to achieve the best classification results.

79 ASTRONOMY AND ASTROPHYSICS↗

Photometric classification of Hyper Suprime-Cam transients using machine learning

Abstract The advancement of technology has resulted in a rapid increase in supernova (SN) discoveries. The Subaru/Hyper Suprime-Cam (HSC) transient survey, conducted from fall 2016 through spring 2017, yielded 1824 SN candidates. This gave rise to the need for fast type classification for spectroscopic follow-up and prompted us to develop a machine learning algorithm using a deep neural network with highway layers. This algorithm is trained by actual observed cadence and filter combinations such that we can directly input the observed data array without any interpretation. We tested our model with a dataset from the LSST classification challenge (Deep Drilling Field). Our classifier scores an area under the curve (AUC) of 0.996 for binary classification (SN Ia or non-SN Ia) and 95.3% accuracy for three-class classification (SN Ia, SN Ibc, or SN II). Application of our binary classification to HSC transient data yields an AUC score of 0.925. With two weeks of HSC data since the first detection, this classifier achieves 78.1% accuracy for binary classification, and the accuracy increases to 84.2% with the full dataset. This paper discusses the potential use of machine learning for SN type classification purposes.

Takahashi, Ichiro↗