Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning for scientific computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Operator inference with roll outs for learning reduced models from scarce and low-quality data

Data-driven modeling has become a key building block in computational science and engineering. However, data that are available in science and engineering are typically scarce, often polluted with noise and affected by measurement errors and other perturbations, which makes learning the dynamics of systems challenging. Here, in this work, we propose to combine data-driven modeling via operator inference with the dynamic training via roll outs of neural ordinary differential equations. Operator inference with roll outs inherits interpretability, scalability, and structure preservation of traditional operator inference while leveraging the dynamic training via roll outs over multiple time steps to increase stability and robustness for learning from low-quality and noisy data. Numerical experiments with data describing shallow water waves and surface quasi-geostrophic dynamics demonstrate that operator inference with roll outs provides predictive models from training trajectories even if data are sampled sparsely in time and polluted with noise of up to 10%.

97 MATHEMATICS AND COMPUTING↗

NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

28 NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning: Preprint

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

New approaches to Bayesian uncertainty quantification for Nuclear Science (Final Technical Report)

Inverse problems play a central role in experimentation and theory/data comparisons for many areas of modern Nuclear Physics (NP) and High-Energy Physics (HEP). Bayes’s Theorem is a powerful tool for solving Inverse Problems, providing conceptually transparent and unbiased constraints on theoretical parameters and their uncertainties (“Bayesian Inference”) and enabling the quantification of agreement or tension between models and data. However, analyses based on Bayesian Inference are often challenging for NP and HEP applications, either because of the large number of parameters in the problem, the high computational cost, or both. We propose a multi-institutional collaboration to develop and deploy novel Bayesian analysis tools that advance the scientific scope of a broad range of current and future NP experiments. This project brings together NP domain scientists working on several high-profile NP projects for which new, high-performance Bayesian Uncertainty Quantification (“Bayesian UQ”) methods are essential to carry out the science, and data scientists who are developing state-of-the-art methods applicable to these problems. The NP projects in this proposal comprise measurements of the mass and fundamental nature of the neutrino; study of the Quark-Gluon Plasma that filled the early universe; and mapping of natural and anthropogenic radiation environments. While these NP projects have very different scientific goals, with datasets and analysis approaches that differ significantly, they share common requirements for improving computationally intensive Bayesian analyses using advanced Machine Learning algorithms and will benefit strongly from a coherent effort to develop general solutions. This proposal brings together these projects and forefront ML-based data science algorithms to develop such general solutions. The methods developed in this project will also be more widely applicable, thereby advancing science in the larger Nuclear Physics portfolio.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Developing and testing capabilities for simulating cases with heterogeneous land/water surfaces in a novel atmospheric large eddy simulation code

Large eddy simulations (LES) are the primary computational tool used to simulate high Reynolds number three-dimensional turbulent flows. In the context of earth system sciences, particularly atmospheric science, LES are uniquely able to resolve the scales of atmospheric motion that are key for building process-level understanding of boundary layer turbulence, atmosphere-surface interaction, clouds, and cloud-aerosol-chemistry interaction, and are a core limited-area modeling capability. Increasing demands are being placed on LES code bases as growing high performance computing resources allow LES to address a wider range of scientific problems. In addition, LES are emerging as a source of high-quality machine learning training data. These demands necessitate an agile and extensible code base that allows the model to quickly adapt to emergent needs. However, LES have largely relied on legacy Fortran code bases that lack flexibility. A new, Python-based LES capability called Predicting INteractions of Aerosol and Clouds in Large Eddy Simulation (PINACLES) has been developed as part of the Department of Energy’s Earth System Model Development (ESMD) program area’s Enabling Aerosol-cloud interactions at Global convection-permitting scalES (EAGLES) project. PINACLES was developed from the ground up with a philosophy of maximizing scientific throughput, by attempting to optimize for both model throughput and software extensibility. The initial development of PINACLES delivered a state-of-the-art idealized LES capability solving the non-hydrostatic anelastic equations of motion with doubly periodic boundary conditions and idealized homogenous surface boundary conditions. Here we provide a final report on the outcomes of a fiscal year 2021 Seed Laboratory Directed Research Project that extended PINACLES in two key ways. First, PINACLES was coupled to a state-of-the-art land surface model enabling it to simulate spatially inhomogeneous land-atmosphere interactions that are known to control key atmospheric processes. Second, the dynamical core of PINACLES was modified to permit non-periodic boundary conditions. This model enhancement enables simulation of realistic cases with boundary conditions prescribed from atmospheric reanalysis and enables nested simulations conducted on a hierarchy of computational domains with increasing resolution. Together, these extensions to PINACLES make it a formidable modeling capability and expand its potential application to diverse components of DOE’s atmospheric science portfolio.

42 ENGINEERING↗

Physics-informed regularization and structure preservation for learning stable reduced models from data with operator inference

Operator inference learns low-dimensional dynamical-system models with polynomial nonlinear terms from trajectories of high-dimensional physical systems (non-intrusive model reduction). Here, this work focuses on the large class of physical systems that can be well described by models with quadratic and cubic nonlinear terms and proposes a regularizer for operator inference that induces a stability bias onto learned models. The proposed regularizer is physics informed in the sense that it penalizes higher-order terms with large norms and so explicitly leverages the polynomial model form that is given by the underlying physics. This means that the proposed approach judiciously learns from data and physical insights combined, rather than from either data or physics alone. Additionally, a formulation of operator inference is proposed that enforces model constraints for preserving structure such as symmetry and definiteness in linear terms. Numerical results demonstrate that models learned with operator inference and the proposed regularizer and structure preservation are accurate and stable even in cases where using no regularization and Tikhonov regularization leads to models that are unstable.

97 MATHEMATICS AND COMPUTING↗

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

When and why PINNs fail to train: A neural tangent kernel perspective

Physics-informed neural networks (PINNs) have lately received great attention thanks to their flexibility in tackling a wide range of forward and inverse problems involving partial differential equations. However, despite their noticeable empirical success, little is known about how such constrained neural networks behave during their training via gradient descent. More importantly, even less is known about why such models sometimes fail to train at all. Here in this work, we aim to investigate these questions through the lens of the Neural Tangent Kernel (NTK); a kernel that captures the behavior of fully-connected neural networks in the infinite width limit during training via gradient descent. Specifically, we derive the NTK of PINNs and prove that, under appropriate conditions, it converges to a deterministic kernel that stays constant during training in the infinite-width limit. This allows us to analyze the training dynamics of PINNs through the lens of their limiting NTK and find a remarkable discrepancy in the convergence rate of the different loss components contributing to the total training error. To address this fundamental pathology, we propose a novel gradient descent algorithm that utilizes the eigenvalues of the NTK to adaptively calibrate the convergence rate of the total training error. Finally, we perform a series of numerical experiments to verify the correctness of our theory and the practical effectiveness of the proposed algorithms.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Dyn$\mathrm{AMO}$: Multi-agent reinforcement learning for dynamic anticipatory mesh optimization with applications to hyperbolic conservation laws

Here we introduce DynAMO, a reinforcement learning paradigm for Dynamic Anticipatory Mesh Optimization. Adaptive mesh refinement is an effective tool for optimizing computational cost and solution accuracy in numerical methods for partial differential equations. However, traditional adaptive mesh refinement approaches for time-dependent problems typically rely only on instantaneous error indicators to guide adaptivity. As a result, standard strategies often require frequent remeshing to maintain accuracy. In the DynAMO approach, multi-agent reinforcement learning is used to discover new local refinement policies that can anticipate and respond to future solution states by producing meshes that deliver more accurate solutions for longer time intervals. By applying DynAMO to discontinuous Galerkin methods for the linear advection and compressible Euler equations in two dimensions, we demonstrate that this new mesh refinement paradigm can outperform conventional threshold-based strategies while also generalizing to different mesh sizes, remeshing and simulation times, and initial conditions.

97 MATHEMATICS AND COMPUTING↗

Bibliometric review and recent advances in total scattering pair distribution function analysis: 21 years in retrospect

Global research activities have been driven by the quest to develop and characterize novel materials for technological advancements. The total scattering pair distribution function (TSPDF) is a powerful and versatile characterization technique for examining the structural details of diverse complex materials including liquid, amorphous, disordered crystalline, and nanostructured materials. Thus, it is critical to keep track of research progress, identify research gaps, and future research directions of the application of the TSPDF technique in materials development and discovery. In this work, a bibliometric analysis of literature regarding the TSPDF technique between 2000 and 2021 was conducted using datasets retrieved from the Web of Science database. The research trends based on publication outputs, research subject distribution, co-authorships among institutions, countries/regions, co-citation of referenced sources, and keyword co-occurrence are evaluated and discussed herein. The impact of the TSPDF technique is projected to increase due to its importance in probing emerging functional materials, and the advances in specialized facilities and instrumentation among the scientific communities engaged with it. Finally, current and emerging research hotspots related to TSPDF technique such as catalysis, computer modeling and simulation, pharmaceutics, machine learning, hydrogen storage, battery materials, and layered structured materials are also identified and discussed.

36 MATERIALS SCIENCE↗

Hybridizing Machine Learning and Physically-based Earth System Models to Improve Prediction of Multivariate Extreme Events (AI Exploration of Wildland Fire Prediction)

Focal Areas: This project responds to two focal areas identified in the DOE Call for AI4ESP White Papers: 1) Predictive modeling through the use of artificial intelligence (AI) techniques, and 2) insights gleaned from complex data using explainable AI and big data analytics. Science Challenge: Large wildland fires (hereafter wildfires) appearing as high-impact compound climate extreme events are closely related to hydroclimate and water cycle extremes that modulate surface fuel supply and combustibility. These compound events have multivariate climatic features (e.g., temperature, precipitation, relative humidity, wind, lightning) and societal drivers (e.g., forest management, land use change, human caused ignitions). Meanwhile, they induce strong feedbacks to the coupled atmosphere, biosphere, and hydrosphere by perturbing regional and global radiation budget as well as ecological, biogeochemical, and water cycles across multiple spatiotemporal scales. The nonlinear interactions between these natural and anthropogenic components of the Earth system are too complex to be completely and adequately represented in today’s Earth system models (ESMs). The inherent stochastic nature of fire activity at all scales further increases the difficulty of its prediction using ESMs that are usually developed from deterministic equations and parameterizations. Besides, concurrence of long-term (decadal to interdecadal) global climate change and fire regime shifts overlapping with short-term (intraseasonal to interannual) variations of regional fire weather and burning activity confound predictability of these compound extreme events. We propose to address the above scientific challenges by using machine learning (ML)-based data-driven modeling techniques to integrate observations and physically-based ESMs’ simulations in a computationally efficient hybrid prediction system. This prediction system is supposed to characterize the wildfire’s sensitivity to climate and exogenous drivers at high resolution (~ 0.25°) on subseasonal to seasonal (S2S) timescales providing improved predictability and explainability. We will use the system to help identify: (1) What are the computational elements of a hybrid system needed to predict compound climate extreme events such as global wildfires? (2) What are the key drivers (either natural or anthropogenic) that modulate short-term variations of multivariate fire weather and burning activity over different regions? How can one take advantage of those driver-response relationships to improve the predictability of large wildfires on S2S time scales? (3) What are the underlying physical mechanisms and sources of improved predictability? Which ML techniques are optimal in revealing and adapting these mechanisms?

54 ENVIRONMENTAL SCIENCES↗

Exploring Classification of Topological Priors With Machine Learning for Feature Extraction

In many scientific endeavors, increasingly abstract representations of data allow for new interpretive methodologies and conceptualization of phenomena. For example, moving from raw imaged pixels to segmented and reconstructed objects allows researchers new insights and means to direct their studies toward relevant areas. Thus, the development of new and improved methods for segmentation remains an active area of research. With advances in machine learning and neural networks, scientists have been focused on employing deep neural networks such as U-Net to obtain pixel-level segmentations, namely, defining associations between pixels and corresponding/referent objects and gathering those objects afterward. Topological analysis, such as the use of the Morse-Smale complex to encode regions of uniform gradient flow behavior, offers an alternative approach: first, create geometric priors, and then apply machine learning to classify. This approach is empirically motivated since phenomena of interest often appear as subsets of topological priors in many applications. Using topological elements not only reduces the learning space but also introduces the ability to use learnable geometries and connectivity to aid the classification of the segmentation target. Here, in this article, we describe an approach to creating learnable topological elements, explore the application of ML techniques to classification tasks in a number of areas, and demonstrate this approach as a viable alternative to pixel-level classification, with similar accuracy, improved execution time, and requiring marginal training data.

97 MATHEMATICS AND COMPUTING↗

A Comprehensive Review of Latent Space Dynamics Identification Algorithms for Intrusive and Non-Intrusive Reduced-Order-Modeling

Numerical solvers of partial differential equations (PDEs) have been widely employed for simulating physical systems. However, the computational cost remains a major bottleneck in various scientific and engineering applications, which has motivated the development of reduced-order models (ROMs). Recently, machine-learning-based ROMs have gained significant popularity and are promising for addressing some limitations of traditional ROM methods, especially for advection dominated systems. In this chapter, we focus on a particular framework known as Latent Space Dynamics Identification (LaSDI), which transforms the high-fidelity data, governed by a PDE, to simpler and low-dimensional latent-space data, governed by ordinary differential equations (ODEs). These ODEs can be learned and subsequently interpolated to make ROM predictions. Each building block of LaSDI can be easily modulated depending on the application, which makes the LaSDI framework highly flexible. In particular, we present strategies to enforce the laws of thermodynamics into LaSDI models (tLaSDI), enhance robustness in the presence of noise through the weak form (WLaSDI), select high-fidelity training data efficiently through active learning (gLaSDI, GPLaSDI), and quantify the ROM prediction uncertainty through Gaussian processes (GPLaSDI). We demonstrate the performance of different LaSDI approaches on Burgers equation, a non-linear heat conduction problem, and a plasma physics problem, showing that LaSDI algorithms can achieve relative errors of less than a few percent and up to thousands of times speed-ups.

Computational Engineering, Finance, and Science (c↗

Peak Prediction Using Multi Layer Perceptron (MLP) for Edge Computing ASICs Targeting Scientific Applications

High data rate detectors play an integral part in scientific research and their development is actively pursued at High Energy Physics (HEP) facilities around the world. Edge Machine Learning (ML) offers the ability to reduce data rates by integrating ML algorithms into Application Specific Integrated Circuits (ASICs) on the front end electronics. In this work, we explore a set of neural network architectures for predicting the peak amplitudes in the detector's sensor response. We have designed and synthesized several MLP based neural networks comparing their inference accuracy, power consumption, and area targeting for minimal latency. The neural networks are synthesized in a commercial 65nm process. The effect of quantizing the network's weights and biases on hardware performance and area is reported. We also conduct design space exploration to compare between design alternatives

47 OTHER INSTRUMENTATION↗

Building the I (Interoperability) of FAIR for performance reproducibility of large-scale composable workflows in RECUP

Abstract-Scientific computing communities increasingly run their experiments using complex data- and compute-intensive workflows that utilize distributed and heterogeneous architectures targeting numerical simulations and machine learning, often executed on the Department of Energy Leadership Computing Facilities (LCFs). We argue that a principled, systematic approach to implementing FAIR principles at scale, including fine-grained metadata extraction and organization, can help with the numerous challenges to performance reproducibility posed by such workflows. We extract workflow patterns, propose a set of tools to manage the entire life cycle of performance metadata, and aggregate them in an HPC-ready framework for reproducibility (RECUP). We describe the challenges in making these tools interoperable, preliminary work, and lessons learned from this experiment.

97 MATHEMATICS AND COMPUTING↗

Novel Proposals for FAIR, Automated, Recommendable, and Robust Workflows

Lightning talks of the Workflows in Support of Large-Scale Science (WORKS) workshop are a venue where the workflow community (researchers, developers, and users) can discuss work in progress, emerging technologies and frameworks, and training and education materials. This paper summarizes the WORKS 2022 lightning talks, which cover five broad topics: data integrity of scientific workflows; a machine learning-based recommendation system; a Python toolkit for running dynamic ensembles of simulations; a cross-platform, high-performance computing utility for processing shell commands; and a meta(data) framework for reproducing hybrid workflows.

Abhinit, Ishan↗

Accelerate microstructure evolution simulation using graph neural networks with adaptive spatiotemporal resolution

Abstract Surrogate models driven by sizeable datasets and scientific machine-learning methods have emerged as an attractive microstructure simulation tool with the potential to deliver predictive microstructure evolution dynamics with huge savings in computational costs. Taking 2D and 3D grain growth simulations as an example, we present a completely overhauled computational framework based on graph neural networks with not only excellent agreement to both the ground truth phase-field methods and theoretical predictions, but enhanced accuracy and efficiency compared to previous works based on convolutional neural networks. These improvements can be attributed to the graph representation, both improved predictive power and a more flexible data structure amenable to adaptive mesh refinement. As the simulated microstructures coarsen, our method can adaptively adopt remeshed grids and larger timesteps to achieve further speedup. The data-to-model pipeline with training procedures together with the source codes are provided.

36 MATERIALS SCIENCE↗