Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data discover”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Gravitational Lenses in UNIONS and Euclid (GLUE). I. A Search for Strong Gravitational Lenses in UNIONS with Subaru, CFHT, and Pan-STARRS Data

We present the results of our pipeline for discovering strong gravitational lenses in the ongoing Ultraviolet Near-Infrared Optical Northern Survey (UNIONS). We successfully train a deep residual convolutional neural network based on CMU-Deeplens architecture, which is designed to detect strong lenses in ground-based imaging surveys. We train on images of real strong lenses and deploy on a sample of 8 million galaxies in areas with full coverage in the g, r, and i filters—the first multiband search for strong gravitational lenses in UNIONS. Following human inspection and grading, we report the discovery of a total of 1346 new strong-lens candidates, of which 146 are Grade A, 199 are Grade B, and 1001 are Grade C. Of these candidates, 283 have lens galaxy spectroscopic redshifts from the Sloan Digital Sky Survey, and an additional 297 have them from the Dark Energy Spectroscopic Instrument Data Release 1. We find 15 of these systems display evidence of both lens and source galaxy redshifts in spectral superposition. We also report the spectroscopic confirmation of seven lensed sources in high-quality systems, all with z > 2.1, using the Keck Near Infrared Echellette Spectrograph and the Gemini Near-Infrared Spectrograph.

Storfer, Christopher J. [University of Hawaii, Hon↗

Linking Spatiotemporal Biological Data to Predict Harmful Algal Blooms

Cyanobacterial Harmful Algal Blooms (cHABs) have significant impacts on an affected region’s economy, ecology, and human health. The blooms can release toxins that kill fish and poison water for people and animals. The global adverse effects of cHABs are exacerbated by the consequences of climate change and increased pollution. Though the phenomena are well documented, scientists’ efforts to mitigate the damage are hampered by insufficient predictive models and incomplete granular knowledge of cHAB community structure. With a goal of leveraging bioinformatics and machine learning tools to better understand and predict cHABs, we are first exploring water sample data sets. Using nearly four thousand samples from the National Center for Biotechnology Information Sequence Read Archive (NCBI-SRA) across 16 years with latitude and longitude embedded in the metadata, we mapped the location of the samples onto a Lake Erie shape file. We combined information about location, date, and community taxa in the NCBI samples to discover factors that determine cHAB features. The data are separated into three distinct zones, with the majority pooled at the southwest end of the lake and occurring in 2017. The samples are rich in biological data; our next steps are to carry out whole genome sequence analysis and use the community profiles as part of our predictive machine learning model.

59 BASIC BIOLOGICAL SCIENCES↗

Physics and chemistry from parsimonious representations: image analysis via invariant variational autoencoders

Electron, optical, and scanning probe microscopy methods are generating ever increasing volume of image data containing information on atomic and mesoscale structures and functionalities. This necessitates the development of the machine learning methods for discovery of physical and chemical phenomena from the data, such as manifestations of symmetry breaking phenomena in electron and scanning tunneling microscopy images, or variability of the nanoparticles. Variational autoencoders (VAEs) are emerging as a powerful paradigm for the unsupervised data analysis, allowing to disentangle the factors of variability and discover optimal parsimonious representation. Here, we summarize recent developments in VAEs, covering the basic principles and intuition behind the VAEs. The invariant VAEs are introduced as an approach to accommodate scale and translation invariances present in imaging data and separate known factors of variations from the ones to be discovered. We further describe the opportunities enabled by the control over VAE architecture, including conditional, semi-supervised, and joint VAEs. Several case studies of VAE applications for toy models and experimental datasets in Scanning Transmission Electron Microscopy are discussed, emphasizing the deep connection between VAE and basic physical principles. Python codes and datasets discussed in this article are available at https://github.com/saimani5/VAE-tutorials and can be used by researchers as an application guide when applying these to their own datasets.

36 MATERIALS SCIENCE↗

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is that the amount of equipment reliability (ER) data being continuously generated are extremely large. These data elements come in different forms: textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) and they provide system engineers with valuable insights and information regarding the discovery of anomalous behaviors or degradation trends, the identification of the possible causes behind such behaviors and trends, and the prediction of their direct consequences. This paper directly targets the generation of knowledge from ER data by putting “data into context”. Here, we employ model-based system engineering (MBSE) models of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by identifying first which elements of the developed MBSE elements they are referring to. This task is much harder for textual data since the information contained in issue or maintenance reports needs to “be understood” by a computational tool. Here we called this process “knowledge extraction” where our methods to extract knowledge from textual data. Lastly, once numeric and textual ER data elements have been processed and “understood”, we discover possible cause-effect relations among them. This is performed by observing if a logical connection through the MBSE models exists, and if there is a temporal relation among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 MATHEMATICS AND COMPUTING↗

Spatio-temporal dynamics of Hendra virus in Australia reveal stable maintenance of diverse viral clades among Pteropus bats

Hendra virus (HeV) was discovered in 1994 in Australia. Limited genomic data have hindered comprehensive understanding of HeV’s evolutionary dynamics. Here, in this work, we recovered 48 HeV genomes from bats and 9 from horses from Australia between 2016 and 2020, revealing four distinct clades. Each clade was distributed over a large spatial area with multiple clades co-circulating within a single bat roost on the same day and over consecutive years. The diversity and temporal stability of co-circulating clades suggest that viral dynamics are driven by episodic shedding of existing lineages maintained at the population level, rather than immune-driven strain-replacement dynamics. HeV isolates of different clades displayed variation in phenotypic properties but minimal antigenic differences. We provide an overview of evolutionary dynamics, phenotypic properties and assessment of countermeasures for HeV, and provide insights into the processes that maintain virus diversity in bats and influence the potential for viral emergence.

genetic variation↗

Relationship between microporous structure and light gas transport through glassy polymeric membranes revealed by molecular simulations

Microporous glassy polymers are attractive materials for gas separation membranes, due to their high permeability and tailorable selectivity, provided physical aging can be delayed. The archetypal microporous glassy polymer, PTMSP, can be blended with a hyper-crosslinked isatin–triptycene porous polymer network (PPN) to delay physical aging. However, while PPN is effective at reducing physical aging, it also affects the permeability of light gases through PTMSP. Molecular dynamics simulations were used here to shed fundamental light on the mechanisms responsible for these effects. Atomistic models are developed that satisfactorily reproduce experimental observations such as matrix density and cavity size distributions for neat PTMSP as well as for PTMSP–PPN blends. Analysis of the simulation results suggests that physical aging is delayed because the PPN inclusions slow down PTMSP relaxation while reducing the connectivity between free volume pockets. To understand how PPN inclusions affect light gas permeability, the atomistic models developed are used to probe CO 2 and CH 4 diffusion and sorption. These simulations are conducted for PTMSP matrices exhibiting varying density, towards reproducing experimental permeability data. Interrogating the simulation trajectories, it is discovered that while CH 4 travels preferentially through the free volume cavities, CO 2 preferentially interacts with the available surfaces, especially in the presence of PPN. These differences help interpret experimental observations. The transport of C 2 H 6 and H 2 S through PTMSP matrices was also investigated. These gases, and in particular H 2 S, were found to absorb within PTMSP, yielding very low diffusion coefficients. The ability to predict differences in diffusion pathways for various natural gas components unveils the possibility of engineering membranes to control permeability and selectivity towards large-scale applications.

Bao Le, Tran Thi [Univ. of Oklahoma, Norman, OK (U↗

Constrained or unconstrained? Neural-network-based equation discovery from data

Throughout many fields, practitioners often rely on differential equations to model systems. Yet, for many applications, the theoretical derivation of such equations and/or the accurate resolution of their solutions may be intractable. Instead, recently developed methods, including those based on parameter estimation, operator subset selection, and neural networks, allow for the data-driven discovery of both ordinary and partial differential equations (PDEs), on a spectrum of interpretability. The success of these strategies is often contingent upon the correct identification of representative equations from noisy observations of state variables and, as importantly and intertwined with that, the mathematical strategies utilized to enforce those equations. Specifically, the latter has been commonly addressed via unconstrained optimization strategies. Representing the PDE as a neural network, we propose to discover the PDE (or the associated operator) by solving a constrained optimization problem and using an intermediate state representation similar to a physics-informed neural network (PINN). The objective function of this constrained optimization problem promotes matching the data, while the constraints require that the discovered PDE is satisfied at a number of spatial collocation points. We present a penalty method and a widely used trust-region barrier method to solve this constrained optimization problem, and we compare these methods on numerical examples. Our results on several example problems demonstrate that the latter constrained method outperforms the penalty method, particularly for higher noise levels or fewer collocation points. This work motivates further exploration into using sophisticated constrained optimization methods in scientific machine learning, as opposed to their commonly used, penalty-method or unconstrained counterparts. For both of these methods, we solve these discovered neural network PDEs with classical methods, such as finite difference methods, as opposed to PINNs-type methods relying on automatic differentiation. Here, we briefly highlight how simultaneously fitting the data while discovering the PDE improves the robustness to noise and other small, yet crucial, implementation details.

Data-driven discovery↗

Grey-Box System Identification of Grid-Forming Inverters

This paper demonstrates the use of grey-box system identification methods for simplifying and understanding the nonlinear power dynamics of grid-forming inverters (GFMs). The power and frequency outputs of complex high-order GFM models are fed into system identification software in order to fit them to a predetermined LTI system and learn system parameters such as (synthetic) inertia and droop constants. The same process is then run for a high-order synchronous generator model, and the outputs are fit to the same set of LTI equations. Simulation of a network of GFM inverters with diverse control architecture is also performed for the same process. The intent is threefold: first, to demonstrate the appropriateness of unified LTI models for describing the power and frequency dynamics of individual resources and connected networks, in order to facilitate analysis of larger heterogeneous networked systems; second, to discover the relationship between internal control parameters of GFMs and their externally observed values; and third, to validate that grey-box data-driven system identification techniques can be a valuable tool to discover the values of important parameters in the absence of explicit vendor models.

analytical models↗

Development of Hydropower Biological Evaluation Toolset (HBET): V2.1.9 Release Notes for HBET

The following release notes reflect changes made to HBET for proposed changes to be released in July 2024. Notes are broken up into three sections: 1) Key Improvements, 2) Bug Fixes, and 3) Data Changes • Key Improvements: primary features added and changes to existing features that affect the user experience. • Bug Fixes: Issues discovered or reported that were fixed in the proposed work to be released. • Data Changes: Any work done on the databases directly or the process to calculate data for the system.

13 HYDRO ENERGY↗

Leveraging 13C-Labeling to Assign Molecular Formulas to Unknown Yeast Metabolites

Mass spectrometry analyses have identified tens of thousands of unknown small molecule-associated peaks in different biological specimens. Notably, even the simplest and best studied organisms like Escherichia coli and Saccharomyces cerevisiae yield thousands of unknown peaks. A key question is how many of these reflect actual novel endogenous metabolites. To explore this, Mahieu and Patti used complete 13 C -labeling in E. coli to credential peaks as biological. This reduced the number of unknowns by more than 90%. Here, we carry out similar uniform 13 C-labeling in the Baker’s yeast S. cerevisiae and two less-studied bioenergy-relevant yeasts Rhodotorula toruloides (lipid producer) and Issatchenkia orientalis (organic acid producer). Identification of unknown metabolite peaks and their molecular formulas is facilitated through software tailored for 13 C labeling data and resulting knowledge of carbon atom count. A classification model evaluates the plausibility of each candidate formula, with peaks lacking plausible candidate formulas unlikely to reflect metabolite molecular ions. This approach prioritizes about one hundred candidate abundant unknown metabolites with logical molecular formulas. Most of these are species-specific rather than conserved across yeasts, and more are found in the nonmodel yeasts than S. cerevisiae. Thus, 13 C-labeling data on unknown metabolites highlights the potential for discovering new metabolites and pathways in nonmodel yeasts.

Carbon↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Discovering the Unknowns: A First Step

This article aims at discovering the unknown variables in the system through data analysis. The main idea is to use the time of data collection as a surrogate variable and try to identify the unknown variables by modeling gradual and sudden changes in the data. We use Gaussian process modeling and a sparse representation of the sudden changes to efficiently estimate the large number of parameters in the proposed statistical model. The method is tested on a realistic dataset generated using a one-dimensional implementation of a Magnetized Liner Inertial Fusion (MagLIF) simulation model, and encouraging results are obtained.

42 ENGINEERING↗

A Pride of Satellites in the Constellation Leo? Discovery of the Leo VI Milky Way Satellite Ultra-faint Dwarf Galaxy with DELVE Early Data Release 3

Abstract We report the discovery and spectroscopic confirmation of an ultra-faint Milky Way satellite in the constellation of Leo. This system was discovered as a spatial overdensity of resolved stars observed with Dark Energy Camera (DECam) data from an early version of the third data release of the DECam Local Volume Exploration (or DELVE) survey. The low luminosity ( M V = − 3.5 6 − 0.37 + 0.47 ; L V = 230 0 − 700 + 1200 L ⊙ ), large size ( R 1 / 2 = 9 0 − 30 + 30 pc), and large heliocentric distance ( D = 11 1 − 6 + 9 kpc) are all consistent with the population of ultra-faint dwarf galaxies (UFDs). Using Keck/DEIMOS observations of the system, we were able to spectroscopically confirm nine member stars, while measuring a tentative mass-to-light ratio of 70 0 − 500 + 1400 M ⊙ / L ⊙ and a nonzero metallicity dispersion of σ [ Fe / H ] = 0.1 9 − 0.11 + 0.14 , further confirming Leo VI’s identity as a UFD. While the system has a highly elliptical shape, ϵ = 0.5 4 − 0.29 + 0.19 , we do not find any conclusive evidence that it is tidally disrupting. Moreover, despite the apparent on-sky proximity of Leo VI to members of the proposed Crater-Leo infall group, its smaller heliocentric distance and inconsistent position in energy–angular momentum space make it unlikely that Leo VI is part of the proposed infall group.

79 ASTRONOMY AND ASTROPHYSICS↗

Physics-Informed Active Learning With Simultaneous Weak-Form Latent Space Dynamics Identification

The parametric greedy latent space dynamics identification (gLaSDI) framework has demonstrated promising potential for accurate and efficient modeling of high-dimensional nonlinear physical systems. However, it remains challenging to handle noisy data. Here, to enhance robustness against noise, we incorporate the weak-form estimation of nonlinear dynamics (WENDy) into gLaSDI. In the proposed weak-form gLaSDI (WgLaSDI) framework, an autoencoder and WENDy are trained simultaneously to discover intrinsic nonlinear latent-space dynamics of high-dimensional data. Compared with the standard sparse identification of nonlinear dynamics (SINDy) employed in gLaSDI, WENDy enables variance reduction and robust latent space discovery, therefore leading to more accurate and efficient reduced-order modeling. Furthermore, the greedy physics-informed active learning in WgLaSDI enables adaptive sampling of optimal training data on the fly for enhanced modeling accuracy. The effectiveness of the proposed framework is demonstrated by modeling various nonlinear dynamical problems, including viscous and inviscid Burgers' equations, time-dependent radial advection, and the Vlasov equation for plasma physics. With data that contains 5%–10% Gaussian white noise, WgLaSDI outperforms gLaSDI by orders of magnitude, achieving 1%–7% relative errors. Compared with the high-fidelity models, WgLaSDI achieves 121 to 1779x speed-up.

97 MATHEMATICS AND COMPUTING↗

Machine Learned Empirical Numerical Integrator from Simulated Data

Recently, a number of state-of-the-art surrogate machine learning (ML) models have been designed for global weather and climate prediction, which have been trained using reanalysis data products. Reanalysis data products are constructed using numerical model simulations that combine numerical integration of partial differential equations and parameterization schemes. These products are typically only archived and made available using coarsened spatial and temporal resolutions. This study explores the impact of the numerical generation methods used to produce the training datasets and the temporal resolution of those datasets on machine learning surrogate models. Using the nonlinear vector autoregression (NVAR) machine as an explainable ML technique, simple dynamical systems are emulated with ML models trained on data produced by three classical numerical integration schemes. NVAR is validated as a skillful ML method, capable of producing accurate predictions and, more importantly, reconstructing both the underlying dynamics and the numerical integration scheme used to generate the training data. However, the machine fails to generalize predictions on unseen test data generated by different numerical integration schemes, despite the underlying dynamical system being the same. This result provides a word of caution for the growing field of machine learning emulation of weather and climate dynamics. Furthermore, we illustrate using NVAR that training on temporally coarsened data may increase the required complexity of ML models and potentially introduce new numerical challenges. Finally, we discover that empirical integration schemes with arbitrary time-stepping sizes can be constructed directly from the data, which implies a potential for the development of empirical numerical integration schemes.

54 ENVIRONMENTAL SCIENCES↗

Broad absorption line quasars in the Dark Energy Spectroscopic Instrument Early Data Release

Broad absorption line (BAL) quasars are characterized by gas clouds that absorb flux at the wavelength of common quasar spectral features, although blueshifted by velocities that can exceed $0.1c$. BAL features are interesting as signatures of significant feedback, yet they can also compromise cosmological studies with quasars by distorting the shape of the most prominent quasar emission lines, impacting redshift accuracy and measurements of the matter density distribution traced by the Lyman $\alpha$ forest. We present a catalogue of BAL quasars discovered in the Dark Energy Spectroscopic Instrument (DESI) survey Early Data Release, which were observed as part of DESI Survey Validation, as well as the first two months of the main survey. We describe our method to automatically identify BAL quasars in DESI data, the quantities we measure for each BAL, and investigate the completeness and purity of this method with mock DESI observations. We mask the wavelengths of the BAL features and re-evaluate each BAL quasar redshift, finding new redshifts which are $243\, {\rm km}\, {\rm s}^{-1}$ smaller on average for the BAL quasar sample. These new, more accurate redshifts are important to obtain the best measurements of quasar clustering, especially at small scales. Finally, we present some spectra of rarer classes of BALs that illustrate the potential of DESI data to identify such populations for further study.

79 ASTRONOMY AND ASTROPHYSICS↗

Determination of the spin and parity of all-charm tetraquarks

The traditional quark model accounts for the existence of baryons, such as protons and neutrons, which consist of three quarks, as well as mesons, composed of a quark–antiquark pair. Only recently has substantial evidence started to accumulate for exotic states composed of four or five quarks and antiquarks. The exact nature of their internal structure remains uncertain. Here we report the first measurement of quantum numbers of the recently discovered family of three all-charm tetraquarks, using data collected by the CMS experiment at the Large Hadron Collider from 2016 to 2018 . The angular analysis techniques developed for the discovery and characterization of the Higgs boson have been applied to the new exotic states. Here we show that the quantum numbers for parity P and charge conjugation C symmetries are found to be +1. The spin J of these exotic states is determined to be consistent with 2ħ, while 0ħ and 1ħ are excluded at 95% and 99% confidence levels, respectively. The J PC = 2 ++ assignment implies particular configurations of constituent spins and orbital angular momenta, which constrain the possible internal structure of these tetraquarks.

Physics↗

Model-Based Approaches to Generate Knowledge from Data in a Plant Reliability Context

One challenge that nuclear power plant system engineers are facing is continuous generation of an extremely large amount of equipment reliability (ER) data. These data elements come in textual (e.g., condition reports) and numeric (e.g., generated by monitoring systems) forms. They provide system engineers with valuable insights and information by discovering anomalous behaviors or degradation trends, identifying possible causes behind such behaviors and trends, and predicting their direct consequences. This paper directly targets the knowledge generation from ER data by putting “data into context.” We employ model-based system engineering (MBSE) of systems and assets to represent and capture their architecture and functional (i.e., cause-effect) relations. ER data elements are processed by first identifying which of the developed MBSE elements they are referring to. This task is harder for textual data since the information contained in issue or maintenance reports needs to be “understood” by a computational tool. We called this process “knowledge extraction” since our methods extract knowledge from textual data. Last, once numeric and textual ER data elements have been processed and “understood,” we discover possible cause-effect relations among them. This is performed by observing whether a logical connection through the MBSE models exists, and if there is a temporal relationship among them. The logic and temporal are the two main ingredients to perform “machine reasoning” from ER data.

97 - MATHEMATICS AND COMPUTING↗