Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model assignment”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

New physics in multi-lepton tau decays

Dark particles with lepton-flavor-violating couplings to the tau lepton can induce rare neutrinoless $τ$ decays with large final state multiplicities. We study models where transitions of the type $τ^\pm\to \ell^\pm\,ϕ$, with $ϕ$ a light new particle, initiate a chain of decays in the dark sector that terminate with decays into electrons, muons, or pions. These decay cascades appear as rare five or even seven-body $τ$ decays with multiple reconstructable resonances. We survey several representative models: kinetically mixed dark photon, gauged $L_i-L_j$ models, and other more exotic charge assignments such as chiral $U(1)'$ extensions of the Standard Model. The main new ingredient is the possibility of flavor violation at very high scales. In these models, a number of channels that have not yet been searched for experimentally, such as $τ\to 5μ$, $τ\to 3μ\,2e$, $τ\to μ\,4e$, and hadronic channels like $τ\to μ\,4π$, typically dominate over the previously-considered signatures such as $τ\to 3μ$. While some of the models, such as the gauged $L_i-L_j$ ones, also contain more challenging channels with missing energy due to decays to neutrinos, they can still be searched for via fully visible channels.

Ema, Yohei [Florida U.]↗

Minimizing thickness variation in monolithic U-10Mo fuel foil and Zr interlayer during hot rolling: A microstructure-based finite element method analysis

Low-enriched uranium alloyed with 10 wt. % molybdenum (U-10Mo) has been identified as a promising alternative to highly enriched uranium fuel for the United States’ high performance research reactors. The monolithic U-10Mo fuel plate consists of a metallic U-10Mo fuel foil with a 25 µm Zr interlayer and a relatively thick cladding of aluminum alloy 6061. The Zr interlayer is typically applied during the hot co-rolling process, and this process dictates the uniformity of the Zr interlayer. Thickness variation observed in the U-10Mo and Zr interlayer has been attributed to several sources: the initial grain size of the U-10Mo castings, can materials, rolling temperature, inhomogeneous molybdenum content, and porosity in the cast U-10Mo. This thickness variation limits the ability to meet the dimensional specification; thus, a better understanding of the factors causing the nonuniform thickness is needed. In this work, we used a novel, microstructure-based finite element method to model the hot rolling process to address these concerns. Grain microstructures in U-10Mo were tessellated and explicitly considered in the finite element model. Each grain was assigned a random material property to mimic the grain strength variations induced by different grain orientations. Simulations were performed using six steel can thicknesses, four grain sizes, and with or without a Zr interlayer to investigate the influences of those variables on the thickness nonuniformity. The simulation results showed that a thinner steel can and finer U-10Mo grain size reduce thickness variations in both the U-10Mo fuel foil and Zr interlayer. The direct findings from the simulations and analysis can be used to optimize the hot rolling schedule, reduce fabrication defects, and meet the dimensional specifications. The proposed microstructure-based finite element model can be also coupled with experimental microstructure characterization data, images, and models to simulate multi-pass hot rolling.

36 MATERIALS SCIENCE↗

A mesoscopic link-transmission-model able to track individual vehicles

Macroscopic traffic flow is a common choice for large-scale traffic simulations. These models do not provide individual-specific metrics as outputs. However, this treatment is necessary in agent-based-models, as in, for example, assigning routes based on personal characteristics. Here, in this paper, we propose an extension of the link-transmission-model, an efficient and yet accurate discretization of the Lighthill-Whitham-Richards (LWR) model, which allow vehicles to be tracked individually while keeping the main features of the underlying model. The extension comprises modifying the link and node models to ensure that the flow between links is always at discrete levels. Therefore, every unit of flow is associated with one individual vehicle moving from its current to its next link. An upper bound of the discretization error is provided. We show that the proposed model resembles its continuous counterpart on lane drop, merge, and diverge cases. In addition, we apply the model into three different networks to validate its applicability in large networks. Finally, we also confirm the parameter transferability between continuous and discrete models and that both can well reproduce field data.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Losses of CO and CO 2 upon collision-activated dissociation of substituted 2-methoxyphenoxides after methyl radical loss

Some deprotonated, substituted 2-methoxyphenols fragment by methyl radical loss followed by competing losses of CO and CO 2 upon collisional activation in linear quadrupole ion trap mass spectrometers. These reactions were examined experimentally and computationally in order to determine their mechanisms. For deprotonated vanillin, the CO loss was found to involve ring contraction, with the highest free energy barrier of 58.0 kcal mol -1 (for the syn rotamer; here, syn refers to the relationship between the aldehyde oxygen atom and the oxygen atom of the methoxy group) for the entire process. The atoms lost in this fragmentation are the oxygen atom that was bound to the first eliminated methyl group and the aromatic carbon atom bound to this oxygen. Examination of carbon-13 labeled vanillin supports these assignments. Examination of several model compounds revealed that this reaction requires the presence of an electron-withdrawing substituent in the para-position relative to the phenol moiety. In contrast, the CO 2 loss from deprotonated vanillin occurs via ring opening followed by re-cyclization and then ring contraction, leading to the loss of CO 2 in a process wherein the highest free energy barrier is 84.5 kcal mol -1 (for the syn isomer). The atoms lost in this fragmentation are the carbon and oxygen atoms from the phenoxide group and the oxygen atom that was bound to the first eliminated methyl group, which is supported by examination of carbon-13 labeled vanillin. Despite the higher total free energy requirement, the CO 2 loss is competitive with CO loss, possibly due to favorable entropy and low activation energy (free energy barrier of 48.9 kcal mol -1 ) for the first reaction step (the analogous value for CO loss is 58.0 kcal mol -1 ). The extent of CO 2 loss is strongly affected by substituents – it either is not observed or is very slow for the other compounds studied here. For example, it is substantially less favorable than CO loss (highest free energy barrier 60.9 kcal mol -1 ) for deprotonated acetovanillone for which the first reaction step for CO 2 loss has a substantially greater free energy barrier (60.6 kcal mol-1) than for vanillin.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Atomic-Level Structure of Mesoporous Hexagonal Boron Nitride Determined by High-Resolution Solid-State Multinuclear Magnetic Resonance Spectroscopy and Density Functional Theory Calculations

Mesoporous hexagonal boron nitride (p-BN) has received significant attention over the last decade as a promising candidate for water cleaning/pollutant removal and hydrogen storage applications. In this work, high-resolution solid-state NMR spectroscopy and plane-wave density-functional theory (DFT) calculations are used to obtain an atomic-level description of p-BN. 1 H– 15 N or 1 H– 14 N heteronuclear (HETCOR) correlation experiments recorded with either conventional NMR at room temperature or dynamic nuclear polarization surface-enhanced spectroscopy (DNP-SENS) at ca. 100 K reveal NB 2 H, NBH 2 , NBH 3 + species residing on the edges of BN sheets. Ultra-high field 35.2 T 11 B NMR spectroscopy was used to resolve 11 B NMR signals from BN 3 , BN 2 O x (OH) 1–x (x = 0–1), BNO x (OH) 2–x (x = 0–2), BO x (OH) 3–x (x = 0–3), and BO x (OH) 4–x – (x = 0–4). Importantly, 2D 11 B dipolar double-quantum–single-quantum homonuclear correlation spectra reveal that many pore/defect sites are composed of boron oxide/hydroxide clusters connected to the BN framework through BN 2 O units. 1D and 2D 11 B{ 15 N} HETCOR NMR experiments, in addition to plane-wave DFT calculations of nine different structural models, further confirm the assignment of all NMR signals. The detailed structure determination of the pore and edge/defect sites within p-BN should further enable the rational design and development of next-generation p-BN-based materials. In addition, the techniques outlined here should be applicable to determine structure within other porous and/or boron-based materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Influence of Polymorphs and Local Defect Structures on NMR Parameters of Graphite Fluorides

In this study, the role of local molecular structure on calculated 13 C and 19 F NMR chemical shifts for graphite fluoride materials was explored by using gauge-including projector augmented wave (GIPAW) computational methods for different periodic crystal polymorphs and density functional theory (DFT) gauge-including atomic orbital (GIAO) computational methods for individual graphite fluoride platelets, i.e., fluorinated graphene (FG). The impact of stacking sequences, d -spacing, and ring conformations on fully fluorinated graphite fluoride structures was investigated. A range of different defects including Stone–Wales, F and C vacancies, void formation, and F inversion were also evaluated using FG structures. These calculations show that distinct chemical shift signatures exist for many of these polymorphs and defects, therefore providing a basis for spectral assignment and development of models describing the mean local CF structure in disordered graphite fluoride materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermal decomposition of neptunyl ammonium nitrate: mechanistic insights and structural characterization of the Np 2 O 5 intermediate phase

Neptunium (Np) possesses a rich and unique chemistry that often diverges from other actinide elements yet remains relatively underexplored compared with the other light actinides. A resurgence of interest in Np has been spurred by the application of 237 Np for plutonium-238 ( 238 Pu) production for use in radioisotope thermoelectric generators (RTGs), necessitating evaluation of Np chemical reactions and materials. The work presented here studied the thermal decomposition of neptunyl ammonium nitrate (NH 4 Np VI O 2 (NO 3 ) 3 ) for synthesis of neptunium dioxide (NpO 2 ), which is the target material used for production of 238 Pu. Additionally, structural characterization of the intermediate solid Np pentoxide (Np 2 O 5 ) was performed. Advanced solid-state characterization techniques, including simultaneous thermal analysis (STA), powder X-ray diffraction (pXRD), Raman spectroscopy, and density functional theory (DFT) modeling have been combined to study the reaction pathways. Analysis revealed that NH 4 Np VI O 2 (NO 3 ) 3 thermally decomposes to a proposed neptunyl nitrate intermediate, followed by Np 2 O 5 and finally NpO 2 , all within the temperature range of 150 °C–600 °C. Further characterization of the pentoxide intermediate provided the first Raman spectra of pure-phase Np 2 O 5 and associated DFT modeling confirmed Raman peak assignments for this phase. These findings provide mechanistic information to advance production of the critical radioisotope 238 Pu and advance the state of knowledge on Np materials chemistry using modern characterization techniques.

Lawson, Kathryn M. [Oak Ridge National Laboratory ↗

Automatic information extraction from childhood cancer pathology reports

The International Classification of Childhood Cancer (ICCC) facilitates the effective classification of a heterogeneous group of cancers in the important pediatric population. However, there has been no development of machine learning models for the ICCC classification. We developed deep learning-based information extraction models from cancer pathology reports based on the ICD-O-3 coding standard. In this article, we describe extending the models to perform ICCC classification. We developed 2 models, ICD-O-3 classification and ICCC recoding (Model 1) and direct ICCC classification (Model 2), and 4 scenarios subject to the training sample size. We evaluated these models with a corpus consisting of 29206 reports with age at diagnosis between 0 and 19 from 6 state cancer registries. Our findings suggest that the direct ICCC classification (Model 2) is substantially better than reusing the ICD-O-3 classification model (Model 1). Applying the uncertainty quantification mechanism to assess the confidence of the algorithm in assigning a code demonstrated that the model achieved a micro-F1 score of 0.987 while abstaining (not sufficiently confident to assign a code) on only 14.8% of ambiguous pathology reports. Our experimental results suggest that the machine learning-based automatic information extraction from childhood cancer pathology reports in the ICCC is a reliable means of supplementing human annotators at state cancer registries by reading and abstracting the majority of the childhood cancer pathology reports accurately and reliably.

60 APPLIED LIFE SCIENCES↗

Generative memory for lifelong machine learning

Techniques are disclosed for training machine learning systems. An input device receives training data comprising pairs of training inputs and training labels. A generative memory assigns training inputs to each archetype task of a plurality of archetype tasks, each archetype task representative of a cluster of related tasks within a task space and assigns a skill to each archetype task. The generative memory generates, from each archetype task, auxiliary data comprising pairs of auxiliary inputs and auxiliary labels. A machine learning system trains a machine learning model to apply a skill assigned to an archetype task to training and auxiliary inputs assigned to the archetype task to obtain output labels corresponding to the training and auxiliary labels associated with the training and auxiliary inputs assigned to the archetype task to enable scalable learning to obtain labels for new tasks for which the machine learning model has not previously been trained.

Nadamuni Raghavan, Aswin↗

Integrated parameter and process learning for hydrologic and biogeochemical modules in Earth System Models

Focus area: Primary focal area #2; secondary focal area #3: Learning about parameters and processes of land surface hydrologic and biogeochemical models in Earth System models by integrating machine learning, physics, and big data. Science challenges: How do we maximally leverage big-data observations to improve hydrobiogeochemical process description and parameterization so that such modules more realistically capture hydrologic and vegetation responses and feedbacks under the future climate? For example, how can we leverage physics, limited observations of vegetation and streamflow to better estimate evapotranspiration, and, relatedly, net primary productivity, especially for drought areas? Vegetation plays a critical role in regional and global water cycles; however, existing vegetation models have failed to predict vegetation response to droughts (McDowell & Xu, 2017) , arctic greening (Keenan & Riley, 2018) , and critical transitions between forest and savanna (Hirota et al., 2011) . These studies suggest that when we build process-based models (PBM) parameterized from regional and global plant traits, we tend to poorly describe plant adaptation and local-scale competition processes. The models and their associated parameters assigned for different regions in the world are not capturing essential heterogeneity in vegetation responses at finer spatial scales. Many parameters of the land surface models control hydrology and vegetation dynamics at the same time. The heterogeneity in vegetation response is a function of (i) plant type, (ii) plant size, (iii) competition and succession, (iv) environmental controls, and (v) local variations due to the unique ecological community that are very difficult to describe (e.g., the size of gaps resulting from fire that facilitated the coexistence of pioneering species). In the demographic models, only factors (i) and (iv) were captured, and plant types were generally described only by leaf phenology and climate zones. With current demographic models, we generally consider more traits to define plant types (i) and calibrate these traits to consider factors (ii), (iii) and (iv); however, it is substantially challenging to scale to regional and global simulations due to trait variations across space (Ali et al., 2016). Moreover, it has been noted that hillslope processes, including ridge-to-valley flow and sunny vs. shady slopes are primary organizers of water, energy, and vegetation (Clark et al., 2015; Fan et al., 2019) . Although gradual improvements in the hydrologic model component in earth system models may reduce this error (at a remarkably slow pace), the long-term, gradual impact of hydrology on plant traits are not well captured. Recent work showed that the hydrologic controls exerted by groundwater and lateral flow are primary regulators of rooting depth (Fan et al., 2017) . Such hydrologic controls have seldom been reflected in vegetation model parameterizations.

54 ENVIRONMENTAL SCIENCES↗

Identifying schools at high-risk for elevated lead in drinking water using only publicly available data

Estimating the risk of lead contamination of schools' drinking water at the State level is a complex, important, and unexplored challenge. Variable water quality among water systems and changes in water chemistry during distribution affect lead dissolution rates from pipes and fittings. In addition, the locations of lead-bearing plumbing materials are uncertain. We tested the capability of six machine learning models to predict the likelihood of lead contamination of drinking water at the schools' taps using only publicly available datasets. The predictive features used in the models correspond to those with a proven correlation to the dominant, but commonly unavailable, factors that govern lead leaching: the presence of lead-bearing plumbing materials and water quality conducive to lead corrosion. By combining water chemistry data from public reports, socioeconomic information from the US census, and spatial features using Geographic Information Systems, we trained and tested models to estimate the likelihood of lead contaminated tap water in over 8,000 schools across California and Massachusetts. Our best-performing model was a Random Forest, with a 10-fold cross validation score of 0.88 for Massachusetts and 0.78 for California using the average Area Under the Receiver Operating Characteristic Curve (ROC AUC) metric. The model was then used to assign a lead leaching risk category to half of the schools across California (the other half was used for training). There was good agreement between the modeled risk categories and the actual lead leaching outcomes for every school; however, the model overestimated the lead leaching risk in up to 17% of the schools. This model is the first of its kind to offer a tool to predict the risk of lead leaching in schools at the State level. Further use of this model can help deploy limited resources more effectively to prevent childhood lead exposure from school drinking water.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

High-dimensional Data-driven Energy optimization for Multi-Modal Transit Agencies (HD-EMMA) (Final Technical Report)

Public bus transit services in the U.S. are responsible for at least 19.7 million metric tons of CO 2 emission annually. Electric vehicles (EVs) can have a much lower environmental impact than comparable internal combustion engine vehicles (ICEVs), especially in urban areas. Unfortunately, EVs are also much more expensive than ICEVs. As a result, many public transit agencies can afford only mixed fleets of transit vehicles, consisting of EVs, hybrids (HEVs), and ICEVs. Transit agencies that operate such mixed fleets of vehicles face a challenging optimization problem: these agencies need to decide which vehicles are assigned to serving which transit trips. Since the advantage of EVs over ICEVs varies depending on the route and time of day (e.g., the benefit of EVs is higher in slower traffic with frequent stops and lower on highways), the assignment can have a significant effect on energy use and, hence, environmental impact. Through this project, we have developed reference data about energy collections and constructed a set of machine learning models that can accurately predict the energy consumption for the whole fleet at the level of each trip. We have used these models to develop a scheduling and assignment strategy that can rotate the different vehicle types across the transit agencies’ routes. The optimization algorithm ensures that the vehicles are matched to trips considering weather patterns, expected congestion, and road gradients to minimize the overall energy usage. We list the key observations from our project for other practitioners below. Details are available in the report, and the list of source code and our publications are included in the appendix. 1. We have demonstrated the feasibility of collecting, merging and analyzing large volumes of high-resolution real-world telemetry data from a mixed vehicle fleet. To mitigate the inherent noise of the recorded GPS points, the team developed an algorithm that filters data and maps the points onto a street. The algorithm considers previous and subsequent location measurements and different characteristics of nearby streets to determine how likely the vehicle travels on them. Then, the team segmented the time series into disjoint contiguous samples based on adjacent road segments and repeated the outlier detection and removal. For each data point, the team added features corresponding to elevation changes within the samples, weather features, such as temperature, and traffic data, such as speed ratio between actual speed and free-flow speed. 2. We have developed two forms of machine learning models that be used to understand and analyze the energy operations of a mixed vehicle transit fleet. The micro prediction model provides estimates of instantaneous energy prediction for all types of buses (diesel, hybrid, and electric). Such a model is important in evaluating the energy impacts of real-time bus operation strategies, but it is challenging due to diversified driving cycles of transit buses. The model can help the drivers understand the impact of their driving behaviors and short-term congestions. The macro prediction models estimate average energy consumption across the whole trip considering the features: distance traveled, various road-type features, elevation change, day of the week, time of day, various weather features (temperature, humidity, etc.), and traffic features (speed ratio and jam factor). 3. We have demonstrated that it is possible to transfer the machine learning models we have developed in this project to other teams and cities by using inductive transfer learning. We also showed that the performance of the macro energy prediction models can be improved using a multi-task learning approach where the learning parameters are shared between the models being developed for different vehicle types. The advantage of this approach is improved learning performance as the models can exploit common spatio-temporal and environmental characteristics. 4. Finally, we have developed trip and vehicle assignment and scheduling algorithms that use the energy prediction models and develop a trip to vehicle type (diesel, electric, hybrid) assignment for the whole operation to reduce overall emissions and cost. We have shown through simulations that the proposed algorithms can save $\$$ 48,910 in energy costs and 175 metric tons of CO 2 emission annually for CARTA.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

MACHINE LEARNING BASED CHEMICAL EXPLOSIVE MODE ANALYSIS (ML-CEMA)

The software consists of a Machine Learning based Chemical Explosive Mode Analysis (ML-CEMA) tool for advanced computational flame diagnostics. CEMA, originally based on eigen-analysis of the local thermochemical system, is capable of identifying reaction fronts and limit phenomena such as auto-ignition and extinction in practical combustion systems, such as internal combustion engines and gas turbine combustors. However, the original CEMA is computationally expensive for large reaction mechanisms that are typically needed to describe fuel chemistry of practical large-hydrocarbon fuels. This novel ML-CEMA tool employs a ML technique to accelerate the eigen-analysis of the basic CEMA approach by orders of magnitude, thus making it suitable for practical fuels. In ML-CEMA,zero-dimensional (0D) reactors and one-dimensional (1D) premixed flames are first used to generate a large number of data points for neural network based ML training. The trained ML model is then used to perform CEMA prediction. This ML-CEMA tool has been demonstrated in canonical 0D and 1D configurations as well as highly-transient three-dimensional spray flames exhibiting multi-mode turbulent combustion, showing promising results. ML-CEMA, as a standalone tool, can be used for computationally-efficient diagnostics of massive datasets generated from both experiments and simulations. For example, based on spatially resolved measurements of a small set of reactive scalars(such as temperature, hydroxyl radical and formaldehyde), ML-CEMA can effectively identify flame fronts and rare events. ML-CEMA also provides a robust online or offline flame feature detection tool. When used for on-the-fly simulations, ML-CEMA further enables zone-adaptive combustion modeling, in which the predicted eigenvalue is used as a robust mode indicator for judicious assignment of locally-valid combustion models. This ML-CEMA based zone-adaptive model can lead to substantial computational cost savings when used for large-scale simulations of multi-mode combustion systems. Third Party Code Web Page to Download Code Web Page Location of Third Party License

Xu, Chao↗

Uncertainty in Thermal Modeling of Spent Nuclear Fuel Casks

Uncertainty is a key metric in computational modeling that must be evaluated for results to have wide ranging applicability. A well characterized uncertainty range is ideal with clear error bars on results that can be presented to stakeholders. In the field of spent fuel cask modeling, this ideal has been historically difficult to achieve in practice because of the computationally intensive nature of the models used and the difficulty assigning reasonable uncertainties to quantities in as-built systems. The work in this report has been conducted to evaluate the overall state of uncertainty and sensitivity in spent fuel cask models and develop methodologies for evaluating these uncertainties. These methodologies must be practical for engineering applications. They should not require excessive computational resources or calendar time to achieve results. In engineering, the model must be on a scale such that it can be changed and adapted throughout a project as new information is discovered and project goals evolve. This report covers three major modeling task areas that provide an overview of the types of sensitivity and uncertainty present in a spent fuel storage and transportation system. Section 3 discusses sensitivity and uncertainty analysis in the effective thermal conductivity model for the fuel region and applies these results to a single assembly model. Section 4 shows sensitivity analysis of a full cask model in the TN-32B and Section 5 demonstrates the overall uncertainty workflow using Coolant Boiling in Rod Arrays – Spent Fuel Storage and STAR-CCM+ developed from the sensitivity work in the preceding sections.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Interpretable machine learning models classify minerals via spectroscopy

Developing methods to identify mineral species confidently and rapidly from Raman spectral analysis is critical to numerous fields. Traditionally, analysis relies on pattern matching the Raman spectrum of an unknown dataset with a supporting library of well-characterized spectral data, which may prove difficult for environmental samples that are poorly crystalline or phase mixtures. Here, we developed interpretable machine learning models that can classify uranium minerals by secondary oxyanion chemistry and other physicochemical properties based solely on Raman spectra. This new ML method produces a mineral profile of physical and chemical properties for an unknown sample and can rapidly classify or identify unknown minerals from Raman data, without the need for an exact pattern match in a spectral library. Training models are validated by 1. Strong correlation of high confidence model regions with published spectroscopic assignments and 2. Correct classification of a mineral not present in training data. Training data are from the Compendium of Uranium Raman and Infrared Experimental Spectra and available crystallographic information files within the open-source Smart Spectral Matching scientific framework. Physically meaningful classifier models can rapidly identify key structural and chemical information about unknown uranium minerals and the overall methodology is broadly applicable for mineral phases.

Machine learning↗

A data-driven perspective on the colours of metal–organic frameworks

Colour is at the core of chemistry and has been fascinating humans since ancient times. It is also a key descriptor of optoelectronic properties of materials and is often used to assess the success of a synthesis. However, predicting the colour of a material based on its structure is challenging. In this work, we leverage subjective and categorical human assignments of colours to build a model that can predict the colour of compounds on a continuous scale. In the process of developing the model, we also uncover inadequacies in current reporting mechanisms. For example, we show that the majority of colour assignments are subject to perceptive spread that would not comply with common printing standards. To remedy this, we suggest and implement an alternative way of reporting colour—and chemical data in general. All data is captured in an objective, and standardised, form in an electronic lab notebook and subsequently automatically exported to a repository in open formats, from where it can be interactively explored by other researchers. We envision this to be key for a data-driven approach to chemical research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Towards AI-assisted neutrino flavor theory design

Particle physics theories, such as those which explain neutrino flavor mixing, arise from a vast landscape of model-building possibilities. A model’s construction typically relies on the intuition of theorists. It also requires considerable effort to identify appropriate symmetry groups, assign field representations, and extract predictions for comparison with experimental data. We develop Autonomous Model Builder (AMBer), a framework in which a reinforcement learning agent interacts with a streamlined physics software pipeline to search these spaces efficiently. AMBer selects symmetry groups, particle content, and group representation assignments to construct models while minimizing the number of free parameters introduced. We validate our approach in well-studied regions of theory space and extend the exploration to a previously unexamined symmetry group. While demonstrated in the context of neutrino flavor theories, this approach of reinforcement learning with physics software feedback may be extended to other theoretical model-building problems in the future.

Baretz, Jason Benjamin↗