Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Mathematics and Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Machine-learned closure of URANS for stably stratified turbulence: connecting physical timescales & data hyperparameters of deep time-series models

Stably stratified turbulence (SST), a model that is representative of the turbulence found in the oceans and atmosphere, is strongly affected by fine balances between forces and becomes more anisotropic in time for decaying scenarios. Moreover, there is a limited understanding of the physical phenomena described by some of the terms in the Unsteady Reynolds-Averaged Navier–Stokes (URANS) equations—used to numerically simulate approximate solutions for such turbulent flows. Rather than attempting to model each term in URANS separately, it is attractive to explore the capability of machine learning (ML) to model groups of terms, i.e. to directly model the force balances. We develop deep time-series ML for closure modeling of the URANS equations applied to SST. We consider decaying SST which are homogeneous and stably stratified by a uniform density gradient, enabling dimensionality reduction. We consider two time-series ML models: long short-term memory and neural ordinary differential equation. Both models perform accurately and are numerically stable in a posteriori (online) tests. Furthermore, we explore the data requirements of the time-series ML models by extracting physically relevant timescales of the complex system. We find that the ratio of the timescales of the minimum information required by the ML models to accurately capture the dynamics of the SST corresponds to the Reynolds number of the flow. The current framework provides the backbone to explore the capability of such models to capture the dynamics of high-dimensional complex dynamical system like SST flows.

97 MATHEMATICS AND COMPUTING↗

Lattice holography on a quantum computer

We explore the potential application of quantum computers to the examination of lattice holography, which extends to the strongly coupled bulk theory regime. With adiabatic evolution, we compute the ground state of a spin system on a ( 2 + 1 )-dimensional hyperbolic lattice, and measure the spin-spin correlation function on the boundary. Notably, we observe that with achievable resources for coming quantum devices, the correlation function demonstrates an approximate scale-invariant behavior, aligning with the pivotal theoretical predictions of the anti–de Sitter/conformal field theory correspondence. Published by the American Physical Society 2024

97 MATHEMATICS AND COMPUTING↗

An AI-Based 3D Bat Movement Tracking System at Wind Energy Facilities Using Multi-Thermal Video Cameras

The poster at the 15th Wind Wildlife Research Meeting discusses how to leverage the potential of real-time thermal-imaging methodologies in quantifying nocturnal bat activities at wind turbines, using 3D computer vision techniques within a deep learning framework. This innovation enables the automatic detection and classification of bats, birds, and insects in thermal-imaging videos captured at wind turbine sites, facilitating efficient and accurate data analysis for enhanced understanding and mitigation of bat-wind turbine interactions.

AI↗

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING↗

Computational Analysis of Hydraulic Efficiency of Michigan DOT Covers J and K

Drainage structures are used in urban street and highway systems to capture stormwater runoff. These structures typically consist of catch basins fitted with grates, inlets, or combination grate/inlet configurations that collect runoff and convey it through buried drainage systems. They are strategically placed within curb-and-gutter systems to enhance public safety by efficiently removing water from roadways and thereby reducing the risk of hydroplaning. The performance of drainage structures is commonly evaluated in terms of hydraulic efficiency, defined as the percentage of flow captured by the basin relative to the total flow reaching the structure. Understanding the hydraulic performance of these structures allows designers to properly space inlets, resulting in cost-effective designs that also help ensure the safety of the traveling public. MDOT uses a variety of drainage structures for runoff capture, as documented in Michigan Department of Transportation (MDOT) Drainage Manual. Many of these structures incorporate sinusoidal-type grates that are not addressed in HEC-22. Because physical modeling of these structures has been limited, further analysis is needed to verify their capture efficiency. Under current MDOT practice, the capture efficiency of these grates is estimated by assuming performance similar to that of a comparably sized reticuline grate described in HEC22. The first phase of this effort, titled Computational Analysis of Hydraulic Efficiency of Michigan DOT Cover C, focused on evaluating the hydraulic performance of MDOT’s Cover C grate. Cover C was selected as the initial test candidate because its sinusoidal pattern is representative of other MDOT grates, while it is typically used in high-volume, higher-speed applications. A similar version, Cover CX, is used on interstate highways but does not include transverse bars for bicycle safety. The current second phase of the study expands this work to evaluate MDOT’s Covers J and K. These grates were selected for additional analysis to further assess the hydraulic performance of MDOT drainage structures that are not directly represented by grate configurations in HEC-22. The results of this phase will build on the findings from the Cover C analysis and support improved understanding of the capture efficiency of MDOT’s standard drainage grates. This report is intended to serve as a companion document to the earlier study, Computational Analysis of Hydraulic Efficiency of Michigan DOT Cover C [3]. The present work applies the same overall CFD-based evaluation approach to MDOT Covers J and K and compares the resulting performance trends with those previously identified for Cover C. In particular, both studies assess on-grade interception efficiency, sag-location hydraulic capacity, and the effects of partial obstruction relative to HEC-22-based design estimates.

42 ENGINEERING↗

Nek5000/RS performance on advanced GPU architectures

The authors explore performance scalability of the open-source thermal-fluids code, NekRS, on the U.S. Department of Energy's leadership computers, Crusher, Frontier, Summit, Perlmutter, and Polaris. Particular attention is given to analyzing performance and time-to-solution at the strong-scale limit for a target efficiency of 80%, which is typical for production runs on the DOE's high-performance computing systems. Several examples of anomalous behavior are also discussed and analyzed.

97 MATHEMATICS AND COMPUTING↗

Digital Twin: Visualizing the Future of the Power Grid [Slides]

The energy systems are generating data at a scale and complexity that increasingly outpaces our ability to analyze it. This talk argues that rigorous, large-scale visualization is a critical tool for meeting that challenge. Drawing on recent work in the Computational Science Center at the National Laboratory of the Rockies, we trace a progression of visualization research culminating in an immersive digital twin of a real campus power grid. Along the way, we examine why visualization matters at all, where widely-used techniques quietly fail, and what becomes analytically possible when you match the display environment to the scale of the problem.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Architecture and performance of Perlmutter's 35 PB ClusterStor E1000 all-flash file system

NERSC's newest system, Perlmutter, features a 35 PB all-flash Lustre file system built on HPE Cray ClusterStor E1000. Here, we present its architecture, early performance figures, and performance considerations unique to this architecture. We demonstrate the performance of E1000 OSSes through low-level Lustre tests that achieve over 90% of the theoretical bandwidth of the SSDs at the OST and LNet levels. We also show end-to-end performance for both traditional dimensions of I/O performance (peak bulk-synchronous bandwidth) and nonoptimal workloads endemic to production computing (small, incoherent I/Os at random offsets) and compare them to NERSC's previous system, Cori, to illustrate that Perlmutter achieves the performance of a burst buffer and the resilience of a scratch file system. Finally, we discuss performance considerations unique to all-flash Lustre and present ways in which users and HPC facilities can adjust their I/O patterns and operations to make optimal use of such architectures.

97 MATHEMATICS AND COMPUTING↗

An end-to-end deep learning method for solving nonlocal Allen–Cahn and Cahn–Hilliard phase-field models

Here, we propose an efficient end-to-end deep learning method for solving nonlocal Allen–Cahn (AC) and Cahn–Hilliard (CH) phase-field models. One motivation for this effort emanates from the fact that discretized partial differential equation-based AC or CH phase-field models result in diffuse interfaces between phases, with the only recourse for remediation is to severely refine the spatial grids in the vicinity of the true moving sharp interface whose width is determined by a grid-independent parameter that is substantially larger than the local grid size. In this work, we introduce non-mass conserving nonlocal AC or CH phase-field models with regular, logarithmic, or obstacle double-well potentials. Because of non-locality, some of these models feature totally sharp interfaces separating phases. The discretization of such models can lead to a transition between phases whose width is only a single grid cell wide. Another motivation is to use deep learning approaches to ameliorate the otherwise high cost of solving discretized nonlocal phase-field models. To this end, loss functions of the customized neural networks are defined using the residual of the fully discrete approximations of the AC or CH models, which results from applying a Fourier collocation method and a temporal semi-implicit approximation. To address the long-range interactions in the models, we tailor the architecture of the neural network by incorporating a nonlocal kernel as an input channel to the neural network model. We then provide the results of extensive computational experiments to illustrate the accuracy, predictive capabilities, and cost reductions of the proposed method.

42 ENGINEERING↗

Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO

Abstract Streambed grain sizes control river hydro‐biogeochemical (HBGC) processes and functions. However, measuring their quantities, distributions, and uncertainties is challenging due to the diversity and heterogeneity of natural streams. This work presents a photo‐driven, artificial intelligence (AI)‐enabled, and theory‐based workflow for extracting the quantities, distributions, and uncertainties of streambed grain sizes from photos. Specifically, we first trained You Only Look Once, an object detection AI, using 11,977 grain labels from 36 photos collected from nine different stream environments. We demonstrated its accuracy with a coefficient of determination of 0.98, a Nash–Sutcliffe efficiency of 0.98, and a mean absolute relative error of 6.65% in predicting the median grain size of 20 ground‐truth photos representing nine typical stream environments. The AI is then used to extract the grain size distributions and determine their characteristic grain sizes, including the 10th, 50th, 60th, and 84th percentiles, for 1,999 photos taken at 66 sites within a watershed in the Northwest US. The results indicate that the 10th, median, 60th, and 84th percentiles of the grain sizes follow log‐normal distributions, with most likely values of 2.49, 6.62, 7.68, and 10.78 cm, respectively. The average uncertainties associated with these values are 9.70%, 7.33%, 9.27%, and 11.11%, respectively. These data allow for the computation of the quantities, distributions, and uncertainties of streambed HBGC parameters, including Manning's coefficient, Darcy‐Weisbach friction factor, top layer interstitial velocity magnitude, and nitrate uptake velocity. Additionally, major sources of uncertainty in grain sizes and their impact on HBGC parameters are examined.

58 GEOSCIENCES↗

Optimal Twirling Depth for Classical Shadows in the Presence of Noise

The classical shadows protocol is an efficient strategy for estimating properties of an unknown state p using a small number of state copies and measurements. In its original form, it involves twirling the state with unitaries from some ensemble and measuring the twirled state in a fixed basis. It was recently shown that for computing local properties, optimal sample complexity (copies of the state required) is remarkably achieved for unitaries drawn from shallow depth circuits composed of local entangling gates, as opposed to purely local (zero depth) or global twirling (infinite depth) ensembles. Here, we consider the sample complexity as a function of the depth of the circuit, in the presence of noise. We find that this noise has important implications for determining the optimal twirling ensemble. Under fairly general conditions, we (i) show that any single-site noise can be accounted for using a depolarizing noise channel with an appropriate damping parameter f, (ii) compute thresholds f th at which optimal twirling reduces to local twirling for Pauli operators, (iii) nth order Renyi entropies (n ≥2), and (iv) provide a meaningful upper bound t max on the optimal circuit depth for any finite noise strength f, which applies to observables and entanglement entropy measurements. In conclusion, these thresholds strongly constrain the search for optimal strategies to implement shadow tomography and are easily tailored to the experimental system at hand.

97 MATHEMATICS AND COMPUTING↗

AI-Enhanced Co-Design for Next-Generation Microelectronics: Innovating Innovation (Workshop Report)

The Artificial Intelligence Enhanced Co-Design for Next Generation Microelectronics virtual workshop was held April 4-5, 2023, and attended by subject matter experts from universities, industry, and national laboratories. This was the third in a series of workshops to motivate the research community to identify and address major challenges facing microelectronics research and production. The 2023 workshop focused on a set of topics from materials to computing algorithms, and included discussions on relevant federal legislation and such as the Creating Helpful Incentives to Produce Semiconductors and Science Act (CHIPS Act) which was signed into law in the summer of 2022. Talks at the workshop included edge computing in radiation environments, new materials for neuromorphic computing, advanced packaging for microelectronics, and new AI techniques. We also received project updates from several of the Department of Energy (DOE) microelectronics co-design projects funded in the fall of 2021, and from three of the Energy Frontier Research Centers (EFRCs) that had been funded in the fall of 2022. The workshop also conducted a set of breakout discussions around the five principal research directions (PRDs) from the 2018 Department of Energy workshop report: 1) define innovative material, device, and architecture requirements driven by applications, algorithms, and software; 2) revolutionize memory and data storage; 3) re-imagine information flow unconstrained by interconnects; 4) redefine computing by leveraging unexploited physical phenomena; 5) reinvent the electricity grid through new materials, devices, and architectures. We tasked each breakout group to consider one primary PRD (and other PRDs as relevant topics arose during discussions) and to address questions such as whether the research community has embraced co-design as a methodology and whether new developments at any level of innovation from materials to programming models requires the research community to reevaluate the PRDs developed back in 2018.

97 MATHEMATICS AND COMPUTING↗

The Essence of Cryptol: A Denotational Cryptol Interpreter in Coq for Foundational Assurances for Quantum Resistant Cryptosystems

Systems of the utmost consequence need a means to establish authenticity of software and data. Cryptosystems implement authentication, but can be vulnerable to cryptographic and implementation attacks. With the threat of quantum cryptographic attacks, “post-quantum” cryptosystems (PQCs) must be henceforth used in these systems. However, the new cryptography needs new ways to, rigorously and machine-checkably, prove systems free of vulnerabilities. We propose a retargetable capability to rapidly instantiate proven correct postquantum cryptosystems through novel proof-carrying synthesis and proof-automation technique, extending those proven successful on existing systems. This capability is crucial to meeting the cryptographic requirements for future high-consequence systems. Since specifications for high consequence cryptography are presently captured in a domain specific language known as Cryptol. While this can enable convenient fully automated reasoning about Cryptol specificaitons and implementations via the Software Analysis Workbench (SAW), Cryptol has expressivity gaps, so that cryptosystems with probabilistic programming features like Falcon cannot be fully expressed in the language. Moreover, SAW’s automation fails for programs and specificaitons with inductive and recursive structure, as in the Sphincs+ PQC. Finally, Cryptol and SAW together represent some 200,000 lines of unverified Haskell, so that the any guarantees about high consequence cryptography are presently contingent on a large, unverified, yet trusted computing base. The first step of the larger project of agile, assured crpytography is therefore to provide a formal, mechanized semantics for Cryptol, so that the specifications expressed by cryptographers in Cryptol can be reasoned about and compiled into performant implementations with a foundational, machine checkable certificate of correctness. This report describes our work on this first step, culminating in the design of a certified denotational interpreter, in Coq, for core Cryptol.

97 MATHEMATICS AND COMPUTING↗

Collaborative Research: Enhancing Laser-Based Ion Sources with High Data Rate Techniques

This collaborative research project focuses on leveraging advanced machine learning techniques to analyze and optimize data from high-repetition-rate laser experiments. The main goal is to apply modern computing hardware, customized data acquisition firmware/software, and machine learning approaches to improve data analysis and experimental control. The project also explores how methodology can be developed on smaller-scale experimental setups and then translated to larger facilities within DOE's LaserNetUS network. With extensive data collection and modeling, the research aims to predict and optimize experimental parameters to enhance performance and efficiency.

47 OTHER INSTRUMENTATION↗

Bounding irrelevant operators in the 3d Gross-Neveu-Yukawa CFTs

We perform a numerical bootstrap study of scalar operators in the critical 3d Gross-Neveu-Yukawa models, a family of conformal field theories containing N Majorana fermions in the fundamental representation of an O(N) global symmetry. We compute rigorous bounds on the scaling dimensions of the next-to-lowest parity-even and parity-odd singlet scalars at N = 2, 4, and 8. All of these dimensions have lower bounds greater than 3, implying that there are only two relevant singlet scalars and placing constraints on the RG flow structure of these theories.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Disk Failure Dataset from the Campaign Storage System

This dataset consists of 1,389 disk (HDD) failure events collected from the Campaign storage system at LANL. The Campaign system supported various compute platforms throughout its lifespan, including Cielo, Fire, Ice, and notably, the Trinity supercomputer. Each recorded event includes its detection timestamp (in ISO 8601 format) and details such as its location within the storage system—rack, enclosure, and drive slot number. The data, spanning from May 4, 2021, to July 25, 2023 (2 years, 2 months, and 22 days), represents failure events from the terminal years of Campaign's operational period, accounting for 26% of its total operational time.

97 MATHEMATICS AND COMPUTING↗

Surrogate modeling of Cellular-Potts agent-based models as a segmentation task using the U-Net neural network architecture

The Cellular-Potts model is a powerful and ubiquitous framework for developing computational models for simulating complex multicellular biological systems. Cellular-Potts models (CPMs) are often computationally expensive due to the explicit modeling of interactions among large numbers of individual model agents and diffusive fields described by partial differential equations (PDEs). In this work, we develop a convolutional neural network (CNN) surrogate model using a U-Net architecture that accounts for periodic boundary conditions. We use this model to accelerate the evaluation of a mechanistic CPM previously used to investigate in vitro vasculogenesis. The surrogate model was trained to predict 100 computational steps ahead (Monte-Carlo steps, MCS), accelerating simulation evaluations by a factor of 562 times compared to single-core CPM code execution on CPU. Over short timescales of up to 3 recursive evaluations, or 300 MCS, our model captures the emergent behaviors demonstrated by the original Cellular-Potts model such as vessel sprouting, extension and anastomosis, and contraction of vascular lacunae. This approach demonstrates the potential for deep learning to serve as a step toward efficient surrogate models for CPM simulations, enabling faster evaluation of computationally expensive CPM simulations of biological processes.

97 MATHEMATICS AND COMPUTING↗