Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

HPC-driven computational reproducibility in numerical relativity codes: a use case study with IllinoisGRMHD

Abstract Reproducibility of results is a cornerstone of the scientific method. Scientific computing encounters two challenges when aiming for this goal. Firstly, reproducibility should not depend on details of the runtime environment, such as the compiler version or computing environment, so results are verifiable by third-parties. Secondly, different versions of software code executed in the same runtime environment should produceconsistent numerical results for physical quantities. In this manuscript, we test the feasibility of reproducing scientific results obtained using theIllinoisGRMHDcode that is part of an open-source community software for simulation in relativistic astrophysics, theEinstein Toolkit. We verify that numerical results of simulating a single isolated neutron star withIllinoisGRMHDcan be reproduced, and compare them to results reported by the code authors in 2015. We use two different supercomputers: Expanse at SDSC, and Stampede2 at TACC. By compiling the source code archived along with the paper on both Expanse and Stampede2, we find thatIllinoisGRMHDreproduces results published in its announcement paper up to errors comparable to round-off level changes in initial data parameters. We also verify that a current version ofIllinoisGRMHDreproduces these results once we account for bug fixes which have occurred since the original publication.

Astronomy & Astrophysics↗

Position Papers for the ASCR Workshop on the Science of Scientific-Software Development and Use

Software is an increasingly important component in the pursuit of scientific discovery. Both its development and use are essential activities for many scientific teams. At the same time, very little scientific study has been conducted to understand, characterize, and improve the development and use of software for science. Computational science teams have diversified over time to include contributions from domain scientists who provide expertise in scientific and engineering disciplines, applied mathematicians and computer scientists who provide optimal algorithms and data structures, and software and data engineers who provide methodologies and tools adapted and adopted from other software domains. These diverse contributions have enabled tremendous advances in the pursuit of scientific discovery, even as models, computer architectures, and software environments have become more complicated. With this increasing diversity, we believe the next opportunity for qualitative improvement comes from applying the scientific method to understanding, characterizing, and improving how scientific software is developed and used. We believe that this pursuit requires expertise from computational scientists themselves, and from the cognitive and social sciences as well as the software engineering research community. As we look to increase the productivity and sustainability of the scientific-software-development-and-use cycle, a more systematic application of the scientific method to understand processes for software development and use will be a valuable tool to guide future work and result in more usable and sustainable software. This workshop will bring together computer scientists, software engineering researchers, computational scientists, applied mathematicians, social scientists, cognitive scientists, and others, to explore how we can conduct such systematic investigations, what can be learned, and how doing so will benefit the scientific enterprise. The workshop will be structured around a set of breakout sessions, with every attendee expected to participate actively in the discussions. Afterward, workshop attendees — from DOE, industry, and academia — will produce a report for ASCR that summarizes the findings of the workshop.

42 ENGINEERING↗

Towards philosophical reasoning with agentic LLMs: Socratic method for scientific assistance

As large language models (LLMs) become central tools in science, improving their reasoning capabilities is critical for meaningful and trustworthy applications. We introduce a Socratic agent for scientific reasoning, implemented through a structured system prompt that guides LLMs via classical principles of inquiry. Unlike typical prompt engineering or retrieval-based methods, our approach leverages definition, analogy, hypothesis elimination, and other Socratic techniques to generate more coherent, critical, and domain-aware responses. We evaluate the agent across diverse scientific domains and benchmark it on the abstraction and reasoning corpus challenge dataset, achieving 97.15% under a fixed prompting protocol and without fine-tuning or external tools. Expert evaluation shows improved reasoning depth, clarity, and adaptability over conventional LLM outputs, suggesting that structured prompting rooted in philosophical reasoning can improve the scientific utility of language models.

LLM reasoning↗

Decision Science for Machine Learning (DeSciML)

The increasing use of machine learning (ML) models to support high-consequence decision making drives a need to increase the rigor of ML-based decision making. Critical problems ranging from climate change to nonproliferation monitoring rely on machine learning for aspects of their analyses. Likewise, future technologies, such as incorporation of data-driven methods into the stockpile surveillance and predictive failure analysis for weapons components, will all rely on decision-making that incorporates the output of machine learning models. In this project, our main focus was the development of decision scientific methods that combine uncertainty estimates for machine learning predictions, with a domain-specific model of error costs. Other focus areas include uncertainty measurement in ML predictions, designing decision rules using multiobjecive optimization, the value of uncertainty reduction, and decision-tailored uncertainty quantification for probability estimates. By laying foundations for rigorous decision making based on the predictions of machine learning models, these approaches are directly relevant to every national security mission that applies, or will apply, machine learning to data, most of which entail some decision context.

97 MATHEMATICS AND COMPUTING↗

Potential Applications of Quantum Computing at Los Alamos National Laboratory, v0.3.0

Since the scientific revolution in the 16th and 17th centuries, the process of scientific discovery has followed an iterative feedback process of observation, hypothesis development and testing with physical experiments, which is widely referred to as the scientific method. This process remained largely unchanged until the middle of the 20th century, when the emergence of digital computers empowered scientist to build and inspect detailed simulations of physical phenomena. Over the last century, computational tools have transformed modern approaches to scientific discovery by enabling fast and affordable hypothesis testing before physical experiments are conducted, shown in Figure 1-1. Some notable examples include: global climate forecasts to understand how the environment may change over decades [130]; modeling the behavior of plasma to design fusion reactors [59]; and understanding the behavior of molecules in biological processes [161, 223].

36 MATERIALS SCIENCE↗

Self-Driving Laboratories for Chemistry and Materials Science

Self-driving laboratories (SDLs) promise an accelerated application of the scientific method. Through the automation of experimental workflows, along with autonomous experimental planning, SDLs hold the potential to greatly accelerate research in chemistry and materials discovery. This review provides an in-depth analysis of the state-of-the-art in SDL technology, its applications across various scientific disciplines, and the potential implications for research and industry. This review additionally provides an overview of the enabling technologies for SDLs, including their hardware, software, and integration with laboratory infrastructure. Most importantly, this review explores the diverse range of scientific domains where SDLs have made significant contributions, from drug discovery and materials science to genomics and chemistry. We provide a comprehensive review of existing real-world examples of SDLs, their different levels of automation, and the challenges and limitations associated with each domain.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Standards, dissemination, and best practices in systems biology

In this study, the reproducibility of scientific research is crucial to the success of the scientific method. Here, we review the current best practices when publishing mechanistic models in systems biology. We recommend, where possible, to use software engineering strategies such as testing, verification, validation, documentation, versioning, iterative development, and continuous integration. In addition, adhering to the Findable, Accessible, Interoperable, and Reusable modeling principles allows other scientists to collaborate and build off of each other’s work. Existing standards such as Systems Biology Markup Language, CellML, or Simulation Experiment Description Markup Language can greatly improve the likelihood that a published model is reproducible, especially if such models are deposited in well-established model repositories. Where models are published in executable programming languages, the source code and their data should be published as open-source in public code repositories together with any documentation and testing code. For complex models, we recommend container-based solutions where any software dependencies and the run-time context can be easily replicated.

59 BASIC BIOLOGICAL SCIENCES↗

Message in a Bottle—An Update to the Golden Record: 1. Objectives and Key Content of the Message

In the first part of this series, we delve into the foundational aspects of “Message in a Bottle (MIAB)” (henceforth referred to as MIAB). This study builds upon the legacy of the Voyager Golden Records, launched aboard Voyager 1 and 2 in 1977, which aimed to communicate with intelligent species beyond our world. These records not only offer a snapshot of Earth and human civilization but also represent our desire to establish contact with advanced alien civilizations. Given the absence of mutually understood signs, symbols, and semiotic conventions, MIAB, like its predecessor, uses scientific methods to design an innovative means of communication that encapsulates the story of humanity. Our goal is to share our collective knowledge, emotions, innovations, and aspirations in a way that provides a universal, yet contextually relevant, understanding of human society, the evolution of life on Earth, and our hopes and concerns for the future. Through this time and space traveling capsule, we also strive to inspire and unify current and future generations to celebrate and safeguard our shared human experience.

99 GENERAL AND MISCELLANEOUS↗

Information and Statistics in Nuclear Experiment and Theory (ISNET)

As with all empirical sciences, nuclear physics operates in the virtuous cycle of the scientific method: observations inspire theoretical models; models lead to new predictions; predictions are tested in experiments; experiments lead to new observations; and so on. Evaluating what we are inferring, and how certain we are of it, is key to this process. These requirements, and a general interest in applying novel statistical, mathematical, and computational techniques, led to the formation of a dedicated research community entitled “Information and Statistics in Nuclear Experiment and Theory (ISNET)” (https://isnet-series.github.io/), which now includes more than 300 members. While the community’s interests lean toward nuclear theory, the unifying theme for this group is the inference of knowledge from data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

PDB‐101: Molecular Explorations through Biology and Medicine

PDB‐101 is an online portal for teachers, students, and the general public to promote exploration of the structural biology of proteins and nucleic acids ( pdb101.rcsb.org ). Learning about the diverse shapes and functions of these biological macromolecules helps to understand all aspects of biomedicine and agriculture, from protein synthesis to health and disease to biological energy. Why PDB‐101? Researchers around the world are studying these molecules at the atomic level. These 3D structures are freely available at the Protein Data Bank (PDB), the central storehouse of biomolecular structures. This website builds introductory materials to help beginners get started in the basics of biomolecular structure and function (“101”, as in an entry level course) as well as resources for extended learning. Since 2011, PDB‐101 has been developed by the RCSB PDB , a global resource for the advancement of research and education in biology and medicine. Along with our Worldwide PDB collaborators, RCSB PDB curates, annotates, and makes publicly available the PDB data deposited by scientists around the globe. The RCSB PDB then provides a window to these data through a rich online resource with powerful searching, reporting, and visualization tools for researchers. This information is then streamlined for students and teachers at PDB‐101. Features include the ongoing Molecule of the Month series, educational materials such as paper models, posters, molecular animations, educational curricula and more. The section “Guide to Understanding PDB Data” is a primer for detailed PDB‐specific information: PDB Data, Visualizing Structures, Reading Coordinate Files, scientific methods for structure determination, and more. PDB‐101 also runs annual Video Challenges for high school students. Participants create short videos that tell molecular stories that connect structural biology and medicine. Previous topics have included HIV/AIDS, diabetes, and antimicrobial resistance. The 2022 challenge will focus on Molecular Mechanisms of Cancer. PDB‐101 activities are evaluated using user surveys, feedback from in‐person activities, and website analytics. In 2020, PDB‐101 hosted >850,000 users and >2.6 million page views.

Zardecki, Christine↗

Examining Nuisance Aerosol Detections in Light of the Origin of the Screening Process (November 2021)

The evolution of philosophy and computations in the International Data Center (IDC) related to aerosol samples have had profound impacts on the number of recorded detections in the network since routine operations began in 2000. Key decisions from policymakers have been the list of triggering radionuclides, the scheme for categorizing these into interest levels 1-5, and an algorithm for determining when an anthropogenic isotope is seen so often that it is no longer interesting, known as the Exponential Weighted Moving Average (EWMA). These are described in the Operations Manual of the IDC. Key parameters that are controlled by the IDC but for which the IDC receives occasional input from policymakers include the constants in EWMA and the peak significance threshold for individual gamma rays, the latter of which directly leads to determination of the presence or absence of a radionuclide in a sample. There are also changes in computations which the IDC makes and informs policy makers about, such as changes in how background is computed, which could also affect the ease of detecting a peak – real or false. Rather than focus on the quantitative changes due to computation changes, this work records some thinking on how isotopes and peak significance levels were chosen, and the resulting detections seen over 18 years during the buildup of the International Monitoring System (IMS). These detections are considered on a global scale to try to determine the relative impact on monitoring, and in some cases, the nature of their existence. Repeated detections of 131I and 133I are the most troublesome, but they are not so frequent to be a major problem for the Verification Regime. These detections could probably be handled adequately using scientific methods currently under development for xenon backgrounds. It is also somewhat problematic that top-level analysis of aerosol backgrounds has not been reported previous to this. The steep increase in the rate of detections after 2016 are a concern, either in the actual backgrounds or from changes in the calculations methods used to generate the Reviewed Radionuclide Report (RRR.) Final conclusions of the authors are that the computational stability of the RRR is very important. With computational stability, changes can be usefully analyzed as being due to changes in radioactivity in Earth’s atmosphere This report is a distillation into text of a talk given in the Radionuclide Experts Group (RNEG) in Vienna during Working Group B (WGB) in February of 2019. This report does not directly contain any IDC data, only summaries by year, or by isotope, or by location. No specific IDC detection by time, location, or isotope is included.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Extended Application of State LiDAR Datasets in Locating Orphaned Wells in Appalachian Region

Location inaccuracies in historical and state oil and gas well databases present a major challenge in locating these orphaned wells. To address this, modern scientific methods such as Light Detection and Ranging (LiDAR), aerial magnetic remote sensing, and digital GIS products have been employed. LiDAR technology uses light to detect surface area changes, providing detailed surface views. This is a workflow to process LiDAR data for use in locating orphaned wells.

Gorantla, Vijaya [NETL Site Support Contractor, Na↗

Observing Arctic Sea Ice

Our understanding of Arctic sea ice and its wide-ranging influence is deeply rooted in observation. Advancing technologies have profoundly improved our ability to observe Arctic sea ice, document its processes and properties, and describe atmosphere-ice-ocean interactions with unprecedented detail. Yet, our progress toward better understanding the Arctic sea ice system is mired by the stark disparities between observations that tend to be siloed by method, scientific discipline, and application. This article presents a review and philosophical design for observing sea ice and accelerating our understanding of the Arctic sea ice system. We give a brief history of Arctic sea ice observations and showcase the 2018 melt season within the context of five observational themes: spatial heterogeneity, temporal variability, cross-disciplinary science, scalability, and retrieval uncertainty. We synthesize buoy data, optical imagery, satellite retrievals, and airborne measurements to demonstrate how disparate data sets can be woven together to transcend observational-scale issues. The results show that there are limitations to interpreting any single data set alone. However, many of these limitations can be surmounted by combining observations that cross spatial and temporal scales. We conclude the article with pathways toward coordinating across observational platforms in order to: (1) optimize the scientific, operational, and community return on observational investments, and (2) facilitate a richer understanding of Arctic sea ice and its role in the climate system.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty quantification in scientific machine learning: Methods, metrics, and comparisons

Neural networks (NNs) are currently changing the computational paradigm on how to combine data with mathematical laws in physics and engineering in a profound way, tackling challenging inverse and ill-posed problems not solvable with traditional methods. However, quantifying errors and uncertainties in NN-based inference is more complicated than in traditional methods. This is because in addition to aleatoric uncertainty associated with noisy data, there is also uncertainty due to limited data, but also due to NN hyperparameters, overparametrization, optimization and sampling errors as well as model misspecification. Although there are some recent works on uncertainty quantification (UQ) in NNs, there is no systematic investigation of suitable methods towards quantifying the total uncertainty effectively and efficiently even for function approximation, and there is even less work on solving partial differential equations and learning operator mappings between infinite-dimensional function spaces using NNs. In this work, we present a comprehensive framework that includes uncertainty modeling, new and existing solution methods, as well as evaluation metrics and post-hoc improvement approaches. Further, to demonstrate the applicability and reliability of our framework, we present an extensive comparative study in which various methods are tested on prototype problems, including problems with mixed input-output data, and stochastic problems in high dimensions. In the Appendix, we include a comprehensive description of all the UQ methods employed. Further, to help facilitate the deployment of UQ in Scientific Machine Learning research and practice, we present and develop in [1] an open-source Python library (github.com/Crunch-UQ4MI/neuraluq), termed NeuralUQ, that is accompanied by an educational tutorial and additional computational experiments.

11 physics-informed neural networks↗

Reliable edge machine learning hardware for scientific applications

Extreme data rate scientific experiments create massive amounts of data that require efficient ML edge processing. This leads to unique validation challenges for VLSI implementations of ML algorithms: enabling bit-accurate functional simulations for performance validation in experimental software frameworks, verifying those ML models are robust under extreme quantization and pruning, and enabling ultra-fine-grained model inspection for efficient fault tolerance. We discuss approaches to developing and validating reliable algorithms at the scientific edge under such strict latency, resource, power, and area requirements in extreme experimental environments. We study metrics for developing robust algorithms, present preliminary results and mitigation strategies, and conclude with an outlook of these and future directions of research towards the longer-term goal of developing autonomous scientific experimentation methods for accelerated scientific discovery.

Baldi, Tommaso↗

Computational Estimation by Scientific Data Mining with Classical Methods to Automate Learning Strategies of Scientists

Experimental results are often plotted as 2-dimensional graphical plots (aka graphs) in scientific domains depicting dependent versus independent variables to aid visual analysis of processes. Repeatedly performing laboratory experiments consumes significant time and resources, motivating the need for computational estimation. The goals are to estimate the graph obtained in an experiment given its input conditions, and to estimate the conditions that would lead to a desired graph. Existing estimation approaches often do not meet accuracy and efficiency needs of targeted applications. We develop a computational estimation approach called AutoDomainMine that integrates clustering and classification over complex scientific data in a framework so as to automate classical learning methods of scientists. Knowledge discovered thereby from a database of existing experiments serves as the basis for estimation. Challenges include preserving domain semantics in clustering, finding matching strategies in classification, striking a good balance between elaboration and conciseness while displaying estimation results based on needs of targeted users, and deriving objective measures to capture subjective user interests. These and other challenges are addressed in this work. The AutoDomainMine approach is used to build a computational estimation system, rigorously evaluated with real data in Materials Science. Our evaluation confirms that AutoDomainMine provides desired accuracy and efficiency in computational estimation. It is extendable to other science and engineering domains as proved by adaptation of its sub-processes within fields such as Bioinformatics and Nanotechnology.

Computer Science↗

Constrained or unconstrained? Neural-network-based equation discovery from data

Throughout many fields, practitioners often rely on differential equations to model systems. Yet, for many applications, the theoretical derivation of such equations and/or the accurate resolution of their solutions may be intractable. Instead, recently developed methods, including those based on parameter estimation, operator subset selection, and neural networks, allow for the data-driven discovery of both ordinary and partial differential equations (PDEs), on a spectrum of interpretability. The success of these strategies is often contingent upon the correct identification of representative equations from noisy observations of state variables and, as importantly and intertwined with that, the mathematical strategies utilized to enforce those equations. Specifically, the latter has been commonly addressed via unconstrained optimization strategies. Representing the PDE as a neural network, we propose to discover the PDE (or the associated operator) by solving a constrained optimization problem and using an intermediate state representation similar to a physics-informed neural network (PINN). The objective function of this constrained optimization problem promotes matching the data, while the constraints require that the discovered PDE is satisfied at a number of spatial collocation points. We present a penalty method and a widely used trust-region barrier method to solve this constrained optimization problem, and we compare these methods on numerical examples. Our results on several example problems demonstrate that the latter constrained method outperforms the penalty method, particularly for higher noise levels or fewer collocation points. This work motivates further exploration into using sophisticated constrained optimization methods in scientific machine learning, as opposed to their commonly used, penalty-method or unconstrained counterparts. For both of these methods, we solve these discovered neural network PDEs with classical methods, such as finite difference methods, as opposed to PINNs-type methods relying on automatic differentiation. Here, we briefly highlight how simultaneously fitting the data while discovering the PDE improves the robustness to noise and other small, yet crucial, implementation details.

Data-driven discovery↗