Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for scientific computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Unsupervised discovery of extreme weather events using universal representations of emergent organization

Spontaneous self-organization is ubiquitous in systems far from thermodynamic equilibrium. While organized structures that emerge dominate transport properties, universal representations that identify and describe these key objects remain elusive. Here, we introduce a theoretically grounded framework for describing emergent organization that, via data-driven algorithms, is constructive in practice. Its building blocks are spacetime lightcones that embody how information propagates across a system through local interactions. We show that predictive equivalence classes of lightcones—local causal states—capture organized behaviors in complex spatiotemporal systems. Employing an unsupervised physics-informed machine learning algorithm and a high-performance computing implementation, we demonstrate automatically discovering organized structures in two real-world domain science problems. We show that local causal states identify vortices and track their power-law decay behavior in two-dimensional fluid turbulence. We then show how to detect and track familiar extreme weather events—hurricanes and atmospheric rivers—and discover other novel structures associated with precipitation extremes in high-resolution climate data at the grid-cell level.

Rupe, Adam [Pacific Northwest National Laboratory ↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

Automated Credibility Assessments of User Features in Scientific Software

Scientific software (SciSoft) is complex, often containing a mixture of production capabilities co-mingled with features under active research and development. Furthermore, SciSoft is often developed over decades by non-computer scientists who may not have a strong background in or prioritize software architecture design, testing, and quality (e.g., test coverage). These conditions lead to difficulty in understanding which software components or functions implement what user-facing features and therefore those features’ software quality pedigree. This lack of understanding poses challenges in assessing readiness and credibility of user features, and often relies on a SciSoft subject matter expert’s (SME) laborious investigation and assertion. This final report of a one-year Computing and Information Sciences Lab Directed Research and Development project presents a general framework for modeling SciSoft architecture as a direct relationship between user features and the software components/functions that implement them. Our approach leverages automated labeling of the SciSoft’s regression test suite and employs machine learning algorithms to construct the architecture model. We demonstrate this framework on the Solid Mechanics component of the SIERRA multi-physics engineering analysis suite developed at Sandia National Laboratories.

97 MATHEMATICS AND COMPUTING↗

Current and future directions in network biology

Network biology is an interdisciplinary field bridging computational and biological sciences that has proved pivotal in advancing the understanding of cellular functions and diseases across biological systems and scales. Although the field has been around for two decades, it remains nascent. It has witnessed rapid evolution, accompanied by emerging challenges. These stem from various factors, notably the growing complexity and volume of data together with the increased diversity of data types describing different tiers of biological organization. We discuss prevailing research directions in network biology, focusing on molecular/cellular networks but also on other biological network types such as biomedical knowledge graphs, patient similarity networks, brain networks, and social/contact networks relevant to disease spread. In more detail, we highlight areas of inference and comparison of biological networks, multimodal data integration and heterogeneous networks, higher-order network analysis, machine learning on networks, and network-based personalized medicine. Following the overview of recent breakthroughs across these five areas, we offer a perspective on future directions of network biology. Additionally, we discuss scientific communities, educational initiatives, and the importance of fostering diversity within the field. This article establishes a roadmap for an immediate and long-term vision for network biology.

59 BASIC BIOLOGICAL SCIENCES↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Machine Learning in Environmental Chemistry: Application to Surface Complexation Modeling

Environmental chemistry – or biogeochemistry – is the scientific discipline typically invoked when examining and quantifying groundwater or surface water contamination, and nutrient cycling in the environment. Over the last three decades, there have been significant advances in mechanistic model development to describe and predict these complex biogeochemical processes. In particular, surface complexation models (SCMs) have been developed to describe the rock/soil surface reactions of metals and radionuclides, and their partitioning between various mobile species in the aqueous phase or immobile species sorbed on solid surfaces. Often represented by a simplified linear isotherm constant – Kd – in reactive transport models, these reactions play a critical role in many environmental science applications; particularly in contamination risk assessments and nuclear waste disposal performance assessments. In the past several decades, efforts by various institutions across the world have focused on developing SCMs based on datasets from laboratory measurements, including the identification of key parameters such as equilibrium constants.

54 ENVIRONMENTAL SCIENCES↗

Energy Optimization of Light and Heavy-Duty Vehicle Cohorts of Mixed Connectivity, Automation and Propulsion System Capabilities via Meshed V2V-V2I and Expanded Data Sharing (Final Scientific and Technical Report)

Vehicle connectivity and automated driving technologies individually have the potential to decrease energy consumption and/or increase safety on light, medium or heavy duty vehicles to varying degrees depending on the traffic infrastructure and specific driving scenarios. Due to advances in sensing, perception and computing power, research and development emphasis in the mobility sector has shifted away from connectivity. Prior research has shown that driving automation with the absence of connectivity can in certain circumstances increase energy consumption. The effectiveness of synergizing connectivity and driving automation technologies is the focus of this work, specifically applied to vehicle cohorts of mixed composition, light and heavy duty, and powertrains ranging from all electric to conventional internal combustion engine. The project team is led by Michigan Technological University (MTU) and partnered with AVL Mobility Technologies Inc. (AVL), Borg Warner (BW), Traffic Technology Services (TTS), American Center for Mobility (ACM) and Navistar (NAV). The main thrusts for the team are to develop a micro-traffic simulation environment with specific VD&PT system attributes and CAV capabilities, 2) field a vehicle test fleet of mixed classification, propulsion and CAV capacity, 3) develop artificial intelligence (AI) and machine learning (ML) based multi-agent optimization methods for various traffic infrastructures, 4) integrate the virtual environment and the optimization methods then deploy the system as a CAV hardware in the loop (HiL) for the vehicle test fleet and 5) conduct closed track and public road testing to validate simulation and demonstrated energy and mobility improvements at multiple scales. For a cohort of mixed vehicles, the team will demonstrate a reduction of energy consumption of 10-50% at intersection, arterial roadway and limited access highway scenarios through connectivity and automation in simulation and at a closed test track. The energy reduction objectives of the project are summarized in Table 1, indicating the infrastructure and over what distances are relevant considered. Single scenario energy reductions are not relevant and thus, the research team took the approach to vary parameters associated with the infrastructure, vehicle cohort composition and dynamic behavior to generate energy consumption distributions for both unconnected and connected scenarios.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fusion RF Modeling Machine Learning (FusionML_RF) v1.0

FusionML_RF consists of multiple codes and trained machine learning (ML) models that perform low-cost output modeling from the Genray-CQL3D. Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. For example, completing a single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time. On the other hand, these ML models achieve ~ms of inference time with high accuracy across the input parameter space. This software collection consists of multiple components. (1) codes that use ML methods and precomputed Genray-CQL3D simulation output to build regression models that enable approximate computations of Genray-CLQ3D outputs from arbitrary but physically meaningful input parameters (surrogate modeling); (2) three trained models created by the team, using a database of 16,000+ GENRAY/CQL3D simulations, to study the performance of ML models for surrogate modeling; (3) codes that load the trained models and simulation data, and then compute mean squared error between the models' predictions and the ground truth of simulation output data. This collection is being made available in conjunction with a scientific publication about the work to promote reusability and provide an artifact of the scientific work.

Bai, Zhe↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Differentiable modelling to unify machine learning and physical models for geosciences

Process-based modelling offers interpretability and physical consistency in many domains of geosciences but struggles to leverage large datasets efficiently. Machine-learning methods, especially deep networks, have strong predictive skills yet are unable to answer specific scientific questions. Here, in this Perspective, we explore differentiable modelling as a pathway to dissolve the perceived barrier between process-based modelling and machine learning in the geosciences and demonstrate its potential with examples from hydrological modelling. ‘Differentiable’ refers to accurately and efficiently calculating gradients with respect to model variables or parameters, enabling the discovery of high-dimensional unknown relationships. Differentiable modelling involves connecting (flexible amounts of) prior physical knowledge to neural networks, pushing the boundary of physics-informed machine learning. It offers better interpretability, generalizability, and extrapolation capabilities than purely data-driven machine learning, achieving a similar level of accuracy while requiring less training data. Additionally, the performance and efficiency of differentiable models scale well with increasing data volumes. Under data-scarce scenarios, differentiable models have outperformed machine-learning models in producing short-term dynamics and decadal-scale trends owing to the imposed physical constraints. Differentiable modelling approaches are primed to enable geoscientists to ask questions, test hypotheses, and discover unrecognized physical relationships. Future work should address computational challenges, reduce uncertainty, and verify the physical significance of outputs.

58 GEOSCIENCES↗

Accelerating science: The usage of commercial clouds in ATLAS Distributed Computing

The ATLAS experiment at CERN is one of the largest scientific machines built to date and will have ever growing computing needs as the Large Hadron Collider collects an increasingly larger volume of data over the next 20 years. ATLAS is conducting R&D projects on Amazon Web Services and Google Cloud as complementary resources for distributed computing, focusing on some of the key features of commercial clouds: lightweight operation, elasticity and availability of multiple chip architectures. The proof of concept phases have concluded with the cloud-native, vendoragnostic integration with the experiment’s data and workload management frameworks. Google Cloud has been used to evaluate elastic batch computing, ramping up ephemeral clusters of up to O(100k) cores to process tasks requiring quick turnaround. Amazon Web Services has been exploited for the successful physics validation of the Athena simulation software on ARM processors. We have also set up an interactive facility for physics analysis allowing endusers to spin up private, on-demand clusters for parallel computing with up to 4 000 cores, or run GPU enabled notebooks and jobs for machine learning applications. The success of the proof of concept phases has led to the extension of the Google Cloud project, where ATLAS will study the total cost of ownership of a production cloud site during 15 months with 10k cores on average, fully integrated with distributed grid computing resources and continue the R&D projects.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The Electron-Ion Collider - A machine that will unlock the secrets of the strongest force in nature!

The computers and smartphones we use every day depend on what we learned about the atom in the last century. All information technology – and much of our economy today – relies on understanding the electromagnetic force between the atomic nucleus and the electrons that orbit it. The science of that force is well understood, but we still know little about the microcosm within the protons and neutrons that make up the atomic nucleus. That’s where Brookhaven National Laboratory (BNL) comes in. Brookhaven National Laboratory (located in Suffolk County, NY, about 60 miles east of midtown Manhattan) was recently chosen as the building site for an Electron-Ion Collider (EIC), a one-of-a-kind nuclear physics research facility. The EIC will be a discovery machine for unlocking the secrets of the “glue” that binds the building blocks of visible matter in the universe. The machine design will take advantage of the existing and highly optimized Relativistic Heavy Ion Collider (RHIC) that’s been operating at Brookhaven Lab since 2000. Beyond sparking scientific discoveries in a new frontier of fundamental physics, the Electron-Ion Collider will trigger technological breakthroughs that have broad-ranging impact on human health and national challenges.

43 PARTICLE ACCELERATORS↗

Hypothesis testing via AI: Generating physically interpretable models of scientific data with machine learning (Full Technical Report)

Deep learning has demonstrated an exceptional ability to solve complex tasks (an engineering success); however, it has done so at the expense of the ability to generate new knowledge (a scientific failure). We propose an alternative framework—entitled Deep Symbolic Regression (DSR)—in which artificial neural networks (NNs) rapidly generate hypotheses about physical relationships among inputs. This framework bypasses the need to interpret an NN altogether, while still leveraging the representational power of deep learning. The resulting models are tractable mathematical expressions, which are inherently and readily human interpretable and can provide insights into underlying physical phenomena. Further, we fold this methodology into the scientific process by allowing the scientist to directly integrate a priori knowledge and beliefs to accelerate learning. We demonstrate this methodology on symbolic regression—the problem of rediscovering underlying expressions describing a dataset—and achieve state-of-the-art performance across a wide variety of symbolic regression problems. Further, we generalize our DSR framework to apply to the more general class of symbolic optimization problems, in which one seeks to optimize a sequence of symbols or “tokens” under a black-box reward function. Examples of other symbolic optimization problems include neural architecture search and computational antibody design. Our generalized tool, Deep Symbolic Optimization (DSO), has been demonstrated on the task of learning symbolic control policies for reinforcement learning environments, and has been adopted as an enabling capability for computational antibody design.

97 MATHEMATICS AND COMPUTING↗

jaxhps: An elliptic PDE solver built with machine learning in mind

Elliptic partial differential equations (PDEs) can model many physical phenomena, such as electrostatics, acoustics, wave propagation, and diffusion. In scientific machine learning settings, a high-throughput PDE solver may be required to generate a training dataset, run in the inner loop of an iterative algorithm, or interface directly with a deep neural network. To provide value to machine learning users, such a PDE solver must be compatible with standard automatic differentiation frameworks, scale efficiently when run on graphics processing units (GPUs), and maintain high accuracy for a large range of input parameters. We have designed the jaxhps package with these use-cases in mind by implementing a highly efficient and accurate solver for elliptic problems with native hardware acceleration and automatic differentiation support.

97 MATHEMATICS AND COMPUTING↗

Artificial Intelligence for Accelerating Nuclear Applications, Science, and Technology

Artificial intelligence (AI) and machine learning (ML) methods have had significant impacts in science and technology in recent years. These methods for generating models from datasets or logic-based algorithms that emulate aspects of human performance can similarly accelerate the fields of nuclear applications, science, and technology toward the IAEA goals of contributing to peace, health, and prosperity. In order to accomplish advances with AI in general and ML in particular across these fields, IAEA can play a significant role by establishing, hosting and curating centralised resources, including databases, adhering to FAIR (findable, accessible, interoperable and reusable) principles and Open Science best practices, providing stewardship of data sharing, supporting training efforts and development of relevant workforces, as well as enabling connections among the scientific, technology, mathematics, AI and ethics communities. Many areas can benefit from the use of AI in the realm of nuclear applications. In human health, these areas include clinical research, epidemiology, nutrition, medical imaging, radiotherapy and education of health professionals. AI-based tools are also being used to facilitate different clinical tasks in imaging, computer-assisted diagnosis in mammography and lung cancer screening programmes, and dose prediction in nuclear medicine procedures. ML methods in particular may also increase the efficiency and accuracy of the analysis of computerised tomography and dual-energy absorptiometry scans for body composition and bone analysis. The application of AI methods to nuclear and related technologies in food and agriculture can lead to significant advances and improved efficiency in the optimisation of agricultural production, food product development, management of supply chains, food safety and food authenticity control. In the water and environmental sector, AI can help inform policies to mitigate the world’s water problems. The application of AI techniques to hydrology and environmental sciences is expected to improve patterns identification and enable model predictions under a changing climate.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Artificial Intelligence and Machine Learning for Bioenergy Research: Opportunities and Challenges

The integration of artificial intelligence and machine learning (AI/ML) with automated experimentation, genomics, biosystems design, and bioprocessing technologies is poised to revolutionize scientific investigation and, particularly, bioenergy research. To identify the opportunities and challenges in this emerging research area, the U.S. Department of Energy’s (DOE) Biological and Environmental Research program (BER) and Bioenergy Technologies Office (BETO) held a joint virtual workshop on AI/ML for Bioenergy Research (AMBER) on August 23–25, 2022. These interests have since been amplified in a September 2022 Executive Order, “Advancing Biotechnology and Biomanufacturing Innovation for a Sustainable, Safe, and Secure U.S. Bioeconomy,” to promote a whole-of government approach to biotechnology development (White House 2022). Approximately 50 scientists with various backgrounds and expertise from academia, industry, and DOE national laboratories met to discuss the opportunities and challenges of AI/ML for bioenergy research. Workshop participants were tasked with assessing the potential for AI/ML and laboratory automation to advance biological understanding and engineering in general. They particularly examined how integrating AI/ML tools with laboratory automation could accelerate biosystems design and optimize biomanufacturing. Discussions included the data and computational infrastructure needed to augment biosystems design applications and the expertise and workforce development efforts urgently required to shift integrated systems toward bioenergy research more broadly. Participants discussed many existing and future applications of AI/ML for biosystems design ranging from enzymes to plants and microbes, microbiomes, and bioprocess development. They also identified three key categories of scientific and technical opportunities and challenges: high-quality data, AI/ML algorithms, and laboratory automation. Several main takeaways emerged from the workshop: 1. Numerous AI/ML and automated experimentation applications exist for a variety of DOE mission needs in energy and the environment; 2. Exemplary research grand challenges for which AI/ML could provide solutions include: building microbes and microbial communities to specifications, developing closed-loop autonomous design and control for biosystems design, and advancing scale-up and automation; 3. Lack of sufficient high-quality, annotated data hinders the development of AI/ML applications; 4. New and improved AI/ML tools are needed, particularly those meeting the specific needs of the BER and BETO research communities; 5. Trade-offs in performance, cost, and reliability exist between deploying commercially available versus building custom-developed instrumentation and software for automated or autonomous experimentation; translation of manual to automated or autonomous methods is often a nontrivial endeavor; 6. Training a new generation of young scientists who can develop and apply AI/ML tools is needed to solve long-standing scientific challenges in bioenergy research. The integration of AI/ML tools and automated experimentation represents a new data-driven research paradigm complementary to the traditional hypothesis-driven research paradigm. This paradigm accelerates design and optimization of biological systems and processes for a variety of DOE mission needs in energy and the environment. The AMBER workshop broadly explored the potential of this new paradigm for bioenergy research, of particular interest to BER and BETO, and identified key challenges and opportunities that DOE can address in the coming years by leveraging its unique capabilities and resources.

59 BASIC BIOLOGICAL SCIENCES↗

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗