Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Hierarchical Bayesian Modeling for Cosmology: Can NPE reliably replace MCMC?

Hierarchical neural posterior estimation has its place Hierarchical Bayesian Modeling (HBM) combined with MCMC algorithms has been shown to provide more robust and accurate inference for real-world phenomena in which nature takes a nested form. However, MCMC-based inference can be computationally expensive, and its performance often suffers for complex posterior geometries. These costs are especially pertinent for HBM. Studies have recently demonstrated the potential for a flexible, expressive, and amortized hierarchical neural posterior estimator (HNPE) built on Normalizing Flows. These studies have mostly been performed on simple datasets, or they focus on a single parameter from each level of the hierarchy. A systematic study analyzing how both hierarchical methods compare for more complex and realistic datasets is necessary before applying HNPE for scientific measurements. Here, we re-explore the theory behind HNPE and conduct comparative numerical experiments of HNPE and MCMC-based HBM methods on real and synthetic data, including strong gravitational lensing simulations. In particular, we use a suite of diagnostics to show trade-offs in terms of accuracy, precision, time to train or sample, reproducibility, and the need for expert domain knowledge. Especially for higher dimensional and complex posteriors, HNPE is expected to drastically improve on time for inference, accuracy, and precision with an upfront training time cost.

Hur, Rachel [Chicago U.] (ORCID:000900089890445X)↗

A General Framework for Error-controlled Unstructured Scientific Data Compression

Data compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the −6 to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets.

Gong, Qian↗

Machine Learning Methods for Connection RTT and Loss Rate Estimation Using MPI Measurements Under Random Losses

Scientific computations are expected to be increasingly distributed across wide-area networks, and Message Passing Interface (MPI) has been shown to scale to support their communications over long distances. Application-level measurements of MPI operations reflect the connection Round-Trip Time (RTT) and loss rate, and machine learningmethods have been previously developed to estimate them under deterministic periodic losses. In this paper, we consider more complex, random losses with uniform, Poisson and Gaussian distributions. We study five disparate machine leaning methods, with linear and non-linear, and smooth and non-smooth properties, to estimate RTT and loss rate over 10 Gbps connections with 0–366 ms RTT. The diversity and complexity of these estimators combined with the randomness of losses and TCP’s non-linear response together rule out the selection of a single best among them; instead, we fuse them to retain their design diversity. Overall, the results show that accurate estimates can be generated at low loss ratesbut become inaccurate at loss rates 10% and higher, thereby illustrating both their strengths and limitations.

Rao, Nageswara S.↗

Ensemble Simulations on Leadership Computing Systems

Scientific productivity can be enhanced through workflow management tools, relieving large High Performance Computing (HPC) system users from the tedious tasks of scheduling and designing the complex computational execution of scientific applications. This paper presents a study on the usage of ensemble workflow tools to accelerate science using the Summit and Frontier supercomputing systems. The research aims to connect science domain simulations using Oak Ridge Leadership Computing Facility (OLCF) supercomputing platforms with ensemble workflow methods in order to accelerate HPC-enabled discovery and boost scientific impact. We present the coupling, porting and optimization of Radical-Cybertools on three applications: Chroma, NAMD and LAMMPS. The tools augment traditional HPC monolithic runs with a pilot scheduler. Lessons-learned are discussed for physics, biology and materials science applications. We discuss intrinsic limitations of coupling and porting ensemble workflow tools to applications that run on large HPC systems. The origins of technical challenges and their solutions developed during the implementation process are discussed. Data management strategies, OLCF’s policies for ensembles, and natively supported workflow tools are also summarized.

Georgiadou, Antigoni [ORNL] (ORCID:000000020977631↗

Benchmark Dose Analysis of DNA Damage Biomarker Responses Provides Compound Potency and Adverse Outcome Pathway Information for the Topoisomerase II Inhibitor Class of Compounds

Genetic toxicology data have traditionally been utilized for hazard identification to provide a binary call for a compound's risk. Recent advances in the scientific field, especially with the development of high‐throughput methods to quantify DNA damage, have influenced a change of approach in genotoxicity assessment. The in vitro MultiFlow® DNA Damage Assay is one such method which multiplexes γH2AX, p53, phospho‐histone H3 biomarkers into a single‐flow cytometric analysis (Bryce et al., [2016]: Environ Mol Mutagen 57:546–558). This assay was used to study human TK6 cells exposed to each of eight topoisomerase II poisons for 4 and 24 hr. Using PROAST v65.5, the Benchmark Dose approach was applied to the resulting flow cytometric datasets. With “compound” serving as covariate, all eight compounds were combined into a single analysis, per time point and endpoint. The resulting 90% confidence intervals, plotted in Log scale, were considered as the potency rank for the eight compounds. The in vitro MultiFlow data showed a maximum confidence interval span of 1Log, which indicates data of good quality. Patterns observed in the compound potency rank were scrutinized by using the expert rule‐based software program Derek Nexus, developed by Lhasa Limited. Compound sub‐classification and structural alerts were considered contributory to the potencies observed for the topoisomerase II poisons studied herein. The Topo II poison Adverse Outcome Pathway was evaluated with MultiFlow endpoints serving as Key Events. The step‐wise approach described herein can be considered as a foundation for risk assessment of compounds within a specific mode of action of interest. Environ. Mol. Mutagen. 2020. © 2020 Wiley Periodicals, Inc.

Wheeldon, Ryan P.↗

RandONets: Shallow networks with random projections for learning linear and nonlinear operators

Deep neural networks have been extensively used for the solution of both the forward and the inverse problem for dynamical systems. However, their implementation necessitates optimizing a high-dimensional space of parameters and hyperparameters. This fact, along with the requirement of substantial computational resources, pose a barrier to achieving high numerical accuracy, but also interpretability. Here, to address the above challenges, we present Random Projection-based Operator Networks (RandONets): shallow networks with random projections and tailor-made numerical analysis methods that learn accurately and fast linear and nonlinear operators. Building on previous works, we prove that RandOnets are universal approximators of linear and nonlinear operators. Due to their simplicity, RandONets provide a one-step transformation of the input space, facilitating interpretability. For the evaluation of their performance, we focus on operators of PDEs. We show, that RandONets outperform by several orders of magnitude, both in terms of numerical approximation accuracy and computational cost, the “vanilla” DeepONets. Hence, we believe that our method will trigger further developments in the field of scientific machine learning, for the development of new ‘’light”schemes that will provide high accuracy while reducing dramatically the computational cost. A MATLAB toolbox for RandONets, including demos, is available on GitHub at https://github.com/GianlucaFabiani/RandONets.

Interpretable machine learning↗

Investigation of local distortion effects on X-ray absorption of ferroelectric perovskites from first principles simulations

Understanding the role of ferroelectric polarization in modulating the electronic and structural properties of crystals is critical for advancing these materials for overcoming various technological and scientific challenges. However, due to difficulties in performing experimental methods with the required resolution, or in interpreting the results of methods therein, the nanoscale morphology and response of these surfaces to external electric fields has not been properly elaborated. Here, in this work, we investigate the effect of ferroelectric polarization and local distortions in a BaTiO 3 perovskite, using two widely used computational approaches which treat the many-body nature of X-ray excitations using different philosophies, namely the many-body, delta-self-consistent-field determinant (mb-ΔSCF) and the Bethe–Salpeter equation (BSE) approaches. We show that in agreement with our experiments, both approaches consistently predict higher excitations of the main peak in the O–K edge for the surface with upward polarization. However, the mb-ΔSCF approach mostly fails to capture the L 2,3 separations at the Ti–L edge, due to the absence of spin–orbit coupling in Kohn–Sham density functional theory (KS-DFT) at the generalized gradient approximation level. On the other hand, and most promising, we show that application of the GW/BSE approach successfully reproduces the experimental XAS, both the relative peak intensities as well as the L 2,3 separations at the Ti–L edges upon ferroelectric switching. Thus simulated XAS is shown to be a powerful method for capturing the nanoscale structure of complex materials, and we underscore the need for many-body perturbation approaches, with explicit consideration of core-hole and multiplet effects, for capturing the essential physics in these systems.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Comparative Analysis of Standard and Advanced USL Methodologies for Nuclear Criticality Safety

The American National Standards Institute/American Nuclear Society national standards 8.1 and 8.24 provide guidance on the requirements and recommendations for establishing confidence in the results of the computerized models used to support operation with fissionable materials. By design, the guidance is not prescriptive, leaving freedom to the analysts to determine how the various sources of uncertainties are to be statistically aggregated. Due to the involved use of statistics entangled with heuristic recipes, the resulting safety margins are often difficult to interpret. Also, these technical margins are augmented by additional administrative margins, which are required to ensure compliance with safety standards or regulations, eliminating the incentive to understand their differences. With the new resurgent wave of advanced nuclear systems, e.g., advanced reactors, fuel cycles, and fuel concepts, focused on economizing operation, there is a strong need to develop a clear understanding of the uncertainties and their consolidation methods to reduce them in manners that can be scientifically defended. In response, the current studies compare the analyses behind four notable methodologies for upper subcriticality limit estimation that have been documented in the nuclear criticality safety literature: the parametric, nonparametric, Whisper, and TSURFER methodologies. Specifically, the work offers a deep dive into the various assumptions of the noted methodologies, their adequacies, and their limitations to provide guidance on developing confidence for the emergent nuclear systems that are expected to be challenged by the scarcity of experimental data. Here, to limit the scope, the current work focuses on the application of these methodologies to criticality safety experiments, where the goal is to calculate a bias, a bias uncertainty, and a tolerance limit for k eff in support of determining an upper subcriticality limit for nuclear criticality safety.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Roadmap on methods and software for electronic structure based simulations in chemistry and materials

This Roadmap article provides a succinct, comprehensive overview of the state of electronic structure methods and software for molecular and materials simulations. Seventeen distinct sections collect insights by 51 leading scientists in the field. Each contribution addresses the status of a particular area, as well as current challenges and anticipated future advances, with a particular eye towards software related aspects and providing key references for further reading. Foundational sections cover density functional theory and its implementation in real-world simulation frameworks, Green's function based many-body perturbation theory, wave-function based and stochastic electronic structure approaches, relativistic effects and semiempirical electronic structure theory approaches. Subsequent sections cover nuclear quantum effects, real-time propagation of the electronic structure, challenges for computational spectroscopy simulations, and exploration of complex potential energy surfaces. The final sections summarize practical aspects, including computational workflows for complex simulation tasks, the impact of current and future high-performance computing architectures, software engineering practices, education and training to maintain and broaden the community, as well as the status of and needs for electronic structure based modeling from the vantage point of industry environments. Overall, the field of electronic structure software and method development continues to unlock immense opportunities for future scientific discovery, based on the growing ability of computations to reveal complex phenomena, processes and properties that are determined by the make-up of matter at the atomic scale, with high precision.

36 MATERIALS SCIENCE↗

Genome‐enabled exploration of microbial ecology and evolution in the sea: a rising tide lifts all boats

Summary As a young bacteriologist just launching my career during the early days of the ‘microbial revolution’ in the 1980s, I was fortunate to participate in some early discoveries, and collaborate in the development of cross‐disciplinary methods now commonly referred to as "metagenomics". My early scientific career focused on applying phylogenetic and genomic approaches to characterize ‘wild’ bacteria, archaea and viruses in their natural habitats, with an emphasis on marine systems. These central interests have not changed very much for me over the past three decades, but knowledge, methodological advances and new theoretical perspectives about the microbial world certainly have. In this invited ‘How we did it’ perspective, I trace some of the trajectories of my lab's collective efforts over the years, including phylogenetic surveys of microbial assemblages in marine plankton and sediments, development of microbial community gene‐ and genome‐enabled surveys, and application of genome‐guided, cultivation‐independent functional characterization of novel enzymes, pathways and their relationships to in situ biogeochemistry. Throughout this short review, I attempt to acknowledge, all the mentors, students, postdocs and collaborators who enabled this research. Inevitably, a brief autobiographical review like this cannot be fully comprehensive, so sincere apologies to any of my great colleagues who are not explicitly mentioned herein. I salute you all as well!

59 BASIC BIOLOGICAL SCIENCES↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Information Content of JWST NIRSpec Transmission Spectra of Warm Neptunes

Warm Neptunes offer a rich opportunity for understanding exo-atmospheric chemistry. With the upcoming James Webb Space Telescope (JWST), there is a need to elucidate the balance between investments in telescope time versus scientific yield. We use the supervised machine-learning method of the random forest to perform an information content (IC) analysis on a 11-parameter model of transmission spectra from the various NIRSpec modes. The three bluest medium-resolution NIRSpec modes (0.7–1.27 μm, 0.97–1.84 μm, 1.66–3.07 μm) are insensitive to the presence of CO. The reddest medium-resolution mode (2.87–5.10 μm) is sensitive to all of the molecules assumed in our model: CO, CO{sub 2}, CH{sub 4}, C{sub 2}H{sub 2}, H{sub 2}O, HCN, and NH{sub 3}. It competes effectively with the three bluest modes on the information encoded on cloud abundance and particle size. It is also competitive with the low-resolution prism mode (0.6–5.3 μm) on the inference of every parameter except for the temperature and ammonia abundance. We recommend astronomers to use the reddest medium-resolution NIRSpec mode for studying the atmospheric chemistry of 800–1200 K warm Neptunes; its corresponding high-resolution counterpart offers diminishing returns. We compare our findings to previous JWST IC analyses that favor the blue orders and suggest that the reliance on chemical equilibrium could lead to biased outcomes if this assumption does not apply. A simple, pressure-independent diagnostic for identifying chemical disequilibrium is proposed based on measuring the abundances of H{sub 2}O, CO, and CO{sub 2}.

79 ASTRONOMY AND ASTROPHYSICS↗

Threat Reduction Research Networks: Fostering Sustainable Collaborations Through Trainings for Genomics for Biosurveillance

Scientific research communities can be represented as heterogeneous or multidimensional networks encompassing multiple types of entities and relationships. These networks might include researchers, institutions, meetings, and publications, connected by relationships like authorship, employment, and attendance. We describe a method for efficiently and flexibly capturing, storing, and extracting information from multidimensional scientific networks using a graph database. The database structure is based on an ontology that captures allowable types of entities and relationships. This allows us to construct a variety of projections of the underlying multidimensional graph through database queries to answer specific research questions. We demonstrate this process through a study of the U.S. Biological Threat Reduction Program (BTRP), which seeks to develop Threat Reduction Networks to build and strengthen a sustainable international community of biosecurity, biosafety, and biosurveillance experts to address shared biological threat reduction challenges. Networks like these create connectional intelligence among researchers and institutions around the world, and are central to the concept of cooperative threat reduction. Our analysis focuses on a series of seven BTRP genome sequencing training workshops, showing how they created a growing network of participants and countries over time, which is also reflected in coauthorship relationships among attendees. By capturing concept and relationship hierarchies, our ontology-based approach allows us to pose general or specific questions about networks within the same framework. This approach can be applied to other research communities or multidimensional social networks to capture, analyze, and visualize different types of interactions and how they change over time.

59 BASIC BIOLOGICAL SCIENCES↗

Why is the winner the best?

International benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multi- center study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), im- age preprocessing (97%), data curation (79%), and post- processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work.

Eisenmann, Matthias↗

Reliable and Efficient Machine Learning (Final Technical Report)

Modern scientific experiments generate massive amounts of data at a pace much faster than humans can manually analyze. While machine learning has revolutionized commercial data analysis (such as recommending movies or recognizing faces), applying these tools to complex scientific discovery is challenging because scientific answers must be precise, interpretable, and adhere to physical laws. The research under this project aims to develop new mathematical tools and computer algorithms specifically designed for scientific applications. Major progress has been made in automatically cleaning and deconstructing messy experimental data, analyzing the visual information of physical phenomena, determining the underlying physical variables, and providing rig orous mathematical analysis of interesting algorithms and concepts widely used in machine learning. This project addressed the critical gap between our ability to generate massive scientific data and our ability to extract interpretable information from it. We established mathematical foundations for Scientific Machine Learning (SciML) aimed at effective data analytics and automated discovery. Our work focused on three core objectives: (1) developing reliable feature extraction methods for dynamic high-dimensional data, (2) establishing mathematical foundations for discovering dynamics via neural networks, and (3) creating rigorous optimization techniques for these models. Key outcomes come from two fronts. On the practical side, they include the development of algorithms that significantly enhance the extraction of signals from field data, as well as the capability to handle situations that exhibit smooth variations or physical stretching due to temperature changes. They also include the creation of an automated framework for discovering fundamental state variables from raw experimental data, demonstrating the ability to identify intrinsic physical dimensions without prior knowledge of the governing laws. On the theoretical front, the research results in theoretical advances in Optimal Transport, a widely used notion in SciML, specifically regarding functions with fixed-size nodal sets, provide sharp bounds relevant to uncertainty quantification. Meanwhile, the outcomes also include the establishment of convergence theories for nonlocal gradient descent methods, enabling robust optimization with noisy data in high-dimensional settings commonly encountered in scientific modeling. The project also helps creating opportunities to train the next generation of researchers, equipping them with the necessary technical skills for today’s workplace and preparing them for future advances.

97 MATHEMATICS AND COMPUTING↗

Solving sparse finite element problems on neuromorphic hardware

The finite element method (FEM) is one of the most important and ubiquitous numerical methods for solving partial differential equations (PDEs) on computers for scientific and engineering discovery. Applying the FEM to larger and more detailed scientific models has driven advances in high-performance computing for decades. Here we demonstrate that scalable spiking neuromorphic hardware can directly implement the FEM by constructing a spiking neural network that solves the large, sparse, linear systems of equations at the core of the FEM. We show that for the Poisson equation, a fundamental PDE in science and engineering, our neural circuit achieves meaningful levels of numerical accuracy and close to ideal scaling on modern, inherently parallel and energy-efficient neuromorphic hardware, specifically Intel’s Loihi 2 neuromorphic platform. We illustrate extensions to irregular mesh geometries in both two and three dimensions as well as other PDEs such as linear elasticity. Our spiking neural network is constructed from a recurrent network model of the brain’s motor cortex and, in contrast to black-box deep artificial neural network-based methods for PDEs, directly translates the well-understood and trusted mathematics of the FEM to a natively spiking neuromorphic algorithm.

Applied mathematics↗

Numerical methods and hypoexponential approximations for gamma distributed delay differential equations

Abstract Gamma distributed delay differential equations (DDEs) arise naturally in many modelling applications. However, appropriate numerical methods for generic gamma distributed DDEs have not previously been implemented. Modellers have therefore resorted to approximating the gamma distribution with an Erlang distribution and using the linear chain technique to derive an equivalent system of ordinary differential equations (ODEs). In this work, we address the lack of appropriate numerical tools for gamma distributed DDEs in two ways. First, we develop a functional continuous Runge–Kutta (FCRK) method to numerically integrate the gamma distributed DDE without resorting to Erlang approximation. We prove the fourth-order convergence of the FCRK method and perform numerical tests to demonstrate the accuracy of the new numerical method. Nevertheless, FCRK methods for infinite delay DDEs are not widely available in existing scientific software packages. As an alternative approach to solving gamma distributed DDEs, we also derive a hypoexponential approximation of the gamma distributed DDE. This hypoexponential approach is a more accurate approximation of the true gamma distributed DDE than the common Erlang approximation but, like the Erlang approximation, can be formulated as a system of ODEs and solved numerically using standard ODE software. Using our FCRK method to provide reference solutions, we show that the common Erlang approximation may produce solutions that are qualitatively different from the underlying gamma distributed DDE. However, the proposed hypoexponential approximations do not have this limitation. Finally, we apply our hypoexponential approximations to perform statistical inference on synthetic epidemiological data to illustrate the utility of the hypoexponential approximation.

97 MATHEMATICS AND COMPUTING↗

Error mitigation, optimization, and extrapolation on a trapped-ion testbed

Current noisy intermediate-scale quantum (NISQ) trapped-ion devices are subject to errors which can significantly impact the accuracy of calculations if left unchecked. A form of error mitigation called zero noise extrapolation (ZNE) can decrease an algorithm’s sensitivity to these errors without increasing the number of required qubits. Here we explore different methods for integrating this error mitigation technique into the Variational Quantum Eigensolver (VQE) algorithm for calculating the ground state of the HeH + molecule at 0.8 Å in the presence of experimental noise. Using the Quantum Scientific Computing Open User Testbed (QSCOUT) trapped-ion device, we test three methods of scaling noise for extrapolation: time stretching the two-qubit gates, scaling the sideband detuning parameter, and inserting two-qubit gate identity operations into the ansatz circuit. We find that time stretching and sideband detuning scaling fail to scale the noise on our particular hardware in a way that can be extrapolated to zero noise. Scaling our noise with global gate identity insertions and extrapolating after variational optimization, we achieve error suppression of 96.8%, resulting in an energy estimate within –0.004 ± 0.04 hartree of the ground state energy. This is an improvement, but still outside the chemical accuracy threshold of 0.0016 hartree. Furthermore, our results show that the efficacy of this error mitigation technique depends on choosing the correct implementation for a given device architecture.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗