Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

Filming movies of attosecond charge migration in single molecules with high harmonic spectroscopy

Electron migration in molecules is the progenitor of chemical reactions and biological functions after light-matter interaction. Following this ultrafast dynamics, however, has been an enduring endeavor. Here we demonstrate that, by using machine learning algorithm to analyze high-order harmonics generated by two-color laser pulses, we are able to retrieve the complex amplitudes and phases of harmonics of single fixed-in-space molecules. These complex dipoles enable us to construct movies of laser-driven electron migration after tunnel ionization of N 2 and CO 2 molecules at time steps of 50 attoseconds. Moreover, the angular dependence of the migration dynamics is fully resolved. By examining the movies, we observe that electron holes do not just migrate along the laser polarization direction, but may swirl around the atom centers. Our result establishes a general scheme for studying ultrafast electron dynamics in molecules, paving a way for further advance in tracing and controlling photochemical reactions by femtosecond lasers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Combining data and theory for derivable scientific discovery with AI-Descartes

Abstract Scientists aim to discover meaningful formulae that accurately describe experimental data. Mathematical models of natural phenomena can be manually created from domain knowledge and fitted to data, or, in contrast, created automatically from large datasets with machine-learning algorithms. The problem of incorporating prior knowledge expressed as constraints on the functional form of a learned model has been studied before, while finding models that are consistent with prior knowledge expressed via general logical axioms is an open problem. We develop a method to enable principled derivations of models of natural phenomena from axiomatic knowledge and experimental data by combining logical reasoning with symbolic regression. We demonstrate these concepts for Kepler’s third law of planetary motion, Einstein’s relativistic time-dilation law, and Langmuir’s theory of adsorption. We show we can discover governing laws from few data points when logical reasoning is used to distinguish between candidate formulae having similar error on the data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-driven predictions of complex organic mixture permeation in polymer membranes

Membrane-based organic solvent separations are rapidly emerging as a promising class of technologies for enhancing the energy efficiency of existing separation and purification systems. Polymeric membranes have shown promise in the fractionation or splitting of complex mixtures of organic molecules such as crude oil. Determining the separation performance of a polymer membrane when challenged with a complex mixture has thus far occurred in an ad hoc manner, and methods to predict the performance based on mixture composition and polymer chemistry are unavailable. Here, we combine physics-informed machine learning algorithms (ML) and mass transport simulations to create an integrated predictive model for the separation of complex mixtures containing up to 400 components via any arbitrary linear polymer membrane. We experimentally demonstrate the effectiveness of the model by predicting the separation of two crude oils within 6-7% of the measurements. Integration of ML predictors of diffusion and sorption properties of molecules with transport simulators enables for the rapid screening of polymer membranes prior to physical experimentation for the separation of complex liquid mixtures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Autonomous closed-loop mechanistic investigation of molecular electrochemistry via automation

Abstract Electrochemical research often requires stringent combinations of experimental parameters that are demanding to manually locate. Recent advances in automated instrumentation and machine-learning algorithms unlock the possibility for accelerated studies of electrochemical fundamentals via high-throughput, online decision-making. Here we report an autonomous electrochemical platform that implements an adaptive, closed-loop workflow for mechanistic investigation of molecular electrochemistry. As a proof-of-concept, this platform autonomously identifies and investigates an EC mechanism, an interfacial electron transfer ( E step) followed by a solution reaction ( C step), for cobalt tetraphenylporphyrin exposed to a library of organohalide electrophiles. The generally applicable workflow accurately discerns the EC mechanism’s presence amid negative controls and outliers, adaptively designs desired experimental conditions, and quantitatively extracts kinetic information of the C step spanning over 7 orders of magnitude, from which mechanistic insights into oxidative addition pathways are gained. This work opens opportunities for autonomous mechanistic discoveries in self-driving electrochemistry laboratories without manual intervention.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Materials structure–property factorization for identification of synergistic phase interactions in complex solar fuels photoanodes

Abstract Properties can be tailored by tuning composition in high-order composition spaces. For spaces with complex phase behavior, modeling the properties as a function of composition and phase distribution remains a formidable challenge. We present materials structure–property factorization (MSPF) as an approach to automate modeling of such data and identify synergistic phase interactions. MSPF is an interpretable machine learning algorithm that couples phase mapping via Deep Reasoning Networks (DRNets) to matrix factorization-based modeling of the representative properties of each phase in a dataset. MSPF is demonstrated for Bi–Cu–V oxide photoanodes for solar fuel generation, which contains 25 different phase combinations and correspondingly exhibits complex composition-structure-photoactivity relationships. Comparing the measured photoactivity to a learned model for non-interacting phases, synergistic phase interactions are identified to guide further photoactivity optimization and understanding. MSPF identifies synergistic interactions of a BiVO 4 -like phase with both Cu 2 V 2 O 7 -like and CuV 2 O 6 -like phases, creating avenues for understanding complex photoelectrocatalysts.

36 MATERIALS SCIENCE↗

Finding predictive models for singlet fission by machine learning

Singlet fission (SF), the conversion of one singlet exciton into two triplet excitons, could significantly enhance solar cell efficiency. Molecular crystals that undergo SF are scarce. Computational exploration may accelerate the discovery of SF materials. However, many-body perturbation theory (MBPT) calculations of the excitonic properties of molecular crystals are impractical for large-scale materials screening. We use the sure-independence-screening-and-sparsifying-operator (SISSO) machine-learning algorithm to generate computationally efficient models that can predict the MBPT thermodynamic driving force for SF for a dataset of 101 polycyclic aromatic hydrocarbons (PAH101). SISSO generates models by iteratively combining physical primary features. The best models are selected by linear regression with cross-validation. The SISSO models successfully predict the SF driving force with errors below 0.2 eV. Based on the cost, accuracy, and classification performance of SISSO models, we propose a hierarchical materials screening workflow. Three potential SF candidates are found in the PAH101 set.

36 MATERIALS SCIENCE↗

Thermodynamic non-ideality and disorder heterogeneity in actinide silicate solid solutions

Non-ideal thermodynamics of solid solutions can greatly impact materials degradation behavior. We have investigated an actinide silicate solid solution system (USiO 4 –ThSiO 4 ), demonstrating that thermodynamic non-ideality follows a distinctive, atomic-scale disordering process, which is usually considered as a random distribution. Neutron total scattering implemented by pair distribution function analysis confirmed a random distribution model for U and Th in first three coordination shells; however, a machine-learning algorithm suggested heterogeneous U and Th clusters at nanoscale (~2 nm). The local disorder and nanosized heterogeneous is an example of the non-ideality of mixing that has an electronic origin. Partial covalency from the U/Th 5f–O 2p hybridization promotes electron transfer during mixing and leads to local polyhedral distortions. The electronic origin accounts for the strong non-ideality in thermodynamic parameters that extends the stability field of the actinide silicates in nature and under typical nuclear waste repository conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Transmon qubit readout fidelity at the threshold for quantum error correction without a quantum-limited amplifier

Abstract High-fidelity and rapid readout of a qubit state is key to quantum computing and communication, and it is a prerequisite for quantum error correction. We present a readout scheme for superconducting qubits that combines two microwave techniques: applying a shelving technique to the qubit that reduces the contribution of decay error during readout, and a two-tone excitation of the readout resonator to distinguish among qubit populations in higher energy levels. Using a machine-learning algorithm to post-process the two-tone measurement results further improves the qubit-state assignment fidelity. We perform single-shot frequency-multiplexed qubit readout, with a 140 ns readout time, and demonstrate 99.5% assignment fidelity for two-state readout and 96.9% for three-state readout–without using a quantum-limited amplifier.

Chen, Liangyu (ORCID:0000000250722445)↗

Widespread underestimation of rain-induced soil carbon emissions from global drylands

Dryland carbon fluxes, particularly those driven by ecosystem respiration, are highly sensitive to water availability and rain pulses. However, the magnitude of rain-induced carbon emissions remains unclear globally. Here we quantify the impact of rain-pulse events on the carbon balance of global drylands and characterize their spatiotemporal controls. Using eddy-covariance observations of carbon, water and energy fluxes from 34 dryland sites worldwide, we produce an inventory of over 1,800 manually identified rain-induced CO2 pulse events. Based on this inventory, a machine learning algorithm is developed to automatically detect rain-induced CO2 pulse events. Our findings show that existing partitioning methods underestimate ecosystem respiration and photosynthesis by up to 30% during rain-pulse events, which annually contribute 16.9 ± 2.8% of ecosystem respiration and 9.6 ± 2.2% of net ecosystem productivity. We show that the carbon loss intensity correlates most strongly with annual productivity, aridity and soil pH. Finally, we identify a universal decay rate of rain-induced CO2 pulses and use it to bias-correct respiration estimates. Our research highlights the importance of rain-induced carbon emissions for the carbon balance of global drylands and suggests that ecosystem models may largely underrepresent the influence of rain pulses on the carbon cycle of drylands.

Nguyen, Ngoc B↗

Supervised enhancer prediction with epigenetic pattern recognition and targeted validation

Enhancers are important non-coding elements, but they have traditionally been hard to characterize experimentally. The development of massively parallel assays allows the characterization of large numbers of enhancers for the first time. Here, we developed a framework using Drosophila STARR-seq to create shape-matching filters based on meta-profiles of epigenetic features. We integrated these features with supervised machine-learning algorithms to predict enhancers. We further demonstrated that our model could be transferred to predict enhancers in mammals. We comprehensively validated the predictions using a combination of in vivo and in vitro approaches, involving transgenic assays in mice and transduction-based reporter assays in human cell lines (153 enhancers in total). The results confirmed that our model can accurately predict enhancers in different species without re-parameterization. Finally, we examined the transcription factor binding patterns at predicted enhancers versus promoters. Here, we demonstrated that these patterns enable the construction of a secondary model that effectively distinguishes enhancers and promoters.

59 BASIC BIOLOGICAL SCIENCES↗

Cric searchable image database as a public platform for conventional pap smear cytology data

Amidst the current health crisis and social distancing, telemedicine has become an important part of mainstream of healthcare, and building and deploying computational tools to support screening more efficiently is an increasing medical priority. The early identification of cervical cancer precursor lesions by Pap smear test can identify candidates for subsequent treatment. However, one of the main challenges is the accuracy of the conventional method, often subject to high rates of false negative. While machine learning has been highlighted to reduce the limitations of the test, the absence of high-quality curated datasets has prevented strategies development to improve cervical cancer screening. The Center for Recognition and Inspection of Cells (CRIC) platform enables the creation of CRIC Cervix collection, currently with 400 images (1,376 × 1,020 pixels) curated from conventional Pap smears, with manual classification of 11,534 cells. This collection has the potential to advance current efforts in training and testing machine learning algorithms for the automation of tasks as part of the cytopathological analysis in the routine work of laboratories.

59 BASIC BIOLOGICAL SCIENCES↗

A Dataset of 3D Structural and Simulated Transport Properties of Complex Porous Media

Physical processes that occur within porous materials have wide-ranging applications including - but not limited to - carbon sequestration, battery technology, membranes, oil and gas, geothermal energy, nuclear waste disposal, water resource management. The equations that describe these physical processes have been studied extensively; however, approximating them numerically requires immense computational resources due to the complex behavior that arises from the geometrically-intricate solid boundary conditions in porous materials. Here, we introduce a new dataset of unprecedented scale and breadth, DRP-372: a catalog of 3D geometries, simulation results, and structural properties of samples hosted on the Digital Rocks Portal. The dataset includes 1736 flow and electrical simulation results on 217 samples, which required more than 500 core years of computation. This data can be used for many purposes, such as constructing empirical models, validating new simulation codes, and developing machine learning algorithms that closely match the extensive purely-physical simulation. This article offers a detailed description of the contents of the dataset including the data collection, simulation schemes, and data validation.

3D images↗

A mapped dataset of surface ocean acidification indicators in large marine ecosystems of the United States

Mapped monthly data products of surface ocean acidification indicators from 1998 to 2022 on a 0.25° by 0.25° spatial grid have been developed for eleven U.S. large marine ecosystems (LMEs). The data products were constructed using observations from the Surface Ocean CO 2 Atlas, co-located surface ocean properties, and two types of machine learning algorithms: Gaussian mixture models to organize LMEs into clusters of similar environmental variability and random forest regressions (RFRs) that were trained and applied within each cluster to spatiotemporally interpolate the observational data. The data products, called RFR-LMEs, have been averaged into regional timeseries to summarize the status of ocean acidification in U.S. coastal waters, showing a domain-wide carbon dioxide partial pressure increase of 1.4 ± 0.4 μatm yr -1 and pH decrease of 0.0014 ± 0.0004 yr -1 . RFR-LMEs have been evaluated via comparisons to discrete shipboard data, fixed timeseries, and other mapped surface ocean carbon chemistry data products. Regionally averaged timeseries of RFR-LME indicators are provided online through the NOAA National Marine Ecosystem Status web portal.

54 ENVIRONMENTAL SCIENCES↗

Image masks of global ship tracks for NASA MODIS data products

Ship tracks, long thin artificial cloud features formed from the pollutants in ship exhaust, are satellite-observable examples of aerosol-cloud interactions (ACI) that can lead to increased cloud albedo and thus increased solar reflectivity, phenomena of interest in solar radiation management. In addition to ship tracks being of interest to meteorologists and policy makers, their observed cloud perturbations provide benchmark evidence of ACI that remain poorly captured by climate models. To broadly analyze the effects of ship tracks, high-resolution satellite imagery data highlighting their presence are required. To support this, we provide a hand labelled dataset to serve as a benchmark for a variety of subsequent analyses. Established from a previous dataset that identified ship track presence using NASA’s MODIS Aqua satellite imager, our first-of-its-kind dataset is comprised of image masks: capturing full ship track regions, including their contours, emission points and dispersive patterns. In total, 300 images, or around 2,500 masked ship tracks, observed under varying conditions are provided, and may facilitate training of machine learning algorithms to automate extraction.

Atmospheric dynamics↗

Machine learning-based microstructure prediction during laser sintering of alumina

Abstract Predicting material’s microstructure under new processing conditions is essential in advanced manufacturing and materials science. This is because the material’s microstructure hugely influences the material’s properties. We demonstrate an elegant machine learning algorithm that faithfully predicts the microstructure under new conditions, without the need of knowing the governing laws. We name this algorithm, RCWGAN-GP, which is regression-based conditional generative adversarial networks with Wasserstein loss function and gradient penalty. This algorithm was trained with experimental SEM micrographs from laser-sintered alumina under various laser powers. The RCWGAN-GP realistically regenerates the SEM micrographs under the trained laser powers. Impressively, it also faithfully predicts the alumina’s microstructure under unexplored laser powers. The predicted microstructure features, including the morphology of the sintered particles and the pores, match the experimental SEM micrographs very well. We further quantitatively examined the prediction accuracy of the RCWGAN-GP. We trained the algorithm with computer-created micrograph datasets of secondary-phase growth governed by the well-known Johnson–Mehl–Avrami (JMA) equation. The RCWGAN-GP accurately regenerates the micrographs at the trained time series, in terms of the grains’ shapes, sizes, and spatial distributions. More importantly, the predicted secondary phase fraction accurately follows the JMA curve.

08 HYDROGEN↗

Property space mapping of Pseudomonas aeruginosa permeability to small molecules

Two membrane cell envelopes act as selective permeability barriers in Gram-negative bacteria, protecting cells against antibiotics and other small molecules. Significant efforts are being directed toward understanding how small molecules permeate these barriers. In this study, we developed an approach to analyze the permeation of compounds into Gram-negative bacteria and applied it to Pseudomonas aeruginosa, an important human pathogen notorious for resistance to multiple antibiotics. The approach uses mass spectrometric measurements of accumulation of a library of structurally diverse compounds in four isogenic strains of P. aeruginosa with varied permeability barriers. We further developed a machine learning algorithm that generates a deterministic classification model with minimal synonymity between the descriptors. This model predicted good permeators into P. aeruginosa with an accuracy of 89% and precision above 58%. The good permeators are broadly distributed in the property space and can be mapped to six distinct regions representing diverse chemical scaffolds. We posit that this approach can be used for more detailed mapping of the property space and for rational design of compounds with high Gram-negative permeability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Observing flow of He II with unsupervised machine learning

Abstract Time dependent observations of point-to-point correlations of the velocity vector field (structure functions) are necessary to model and understand fluid flow around complex objects. Using thermal gradients, we observed fluid flow by recording fluorescence of $${\text{He}}_{2}^{*}$$ He 2 ∗ excimers produced by neutron capture throughout a ~ cm 3 volume. Because the photon emitted by an excited excimer is unlikely to be recorded by the camera, the techniques of particle tracking (PTV) and particle imaging (PIV) velocimetry cannot be applied to extract information from the fluorescence of individual excimers. Therefore, we applied an unsupervised machine learning algorithm to identify light from ensembles of excimers (clusters) and then tracked the centroids of the clusters using a particle displacement determination algorithm developed for PTV.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

MArVD2: a machine learning enhanced tool to discriminate between archaeal and bacterial viruses in viral datasets

Abstract Our knowledge of viral sequence space has exploded with advancing sequencing technologies and large-scale sampling and analytical efforts. Though archaea are important and abundant prokaryotes in many systems, our knowledge of archaeal viruses outside of extreme environments is limited. This largely stems from the lack of a robust, high-throughput, and systematic way to distinguish between bacterial and archaeal viruses in datasets of curated viruses. Here we upgrade our prior text-based tool (MArVD) via training and testing a random forest machine learning algorithm against a newly curated dataset of archaeal viruses. After optimization, MArVD2 presented a significant improvement over its predecessor in terms of scalability, usability, and flexibility, and will allow user-defined custom training datasets as archaeal virus discovery progresses. Benchmarking showed that a model trained with viral sequences from the hypersaline, marine, and hot spring environments correctly classified 85% of the archaeal viruses with a false detection rate below 2% using a random forest prediction threshold of 80% in a separate benchmarking dataset from the same habitats.

Vik, Dean (ORCID:000000027546899X)↗