Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “CA Training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Linking the subseasonal variability of the East Asia winter monsoon and the Madden-Julian Oscillation through wave disturbances along the subtropical jet

Despite an urgent demand for reliable subseasonal-to-seasonal (S2S) predictions to guide disaster preparedness, our current climate models show limited S2S prediction skill, particularly for precipitation, due to an inadequate understanding of the key processes that drive regional S2S variability. Here we demonstrate that the leading subseasonal variability mode of precipitation over the East Asian Winter Monsoon (EAWM) region is not only closely tied to the activity of the Madden-Julian Oscillation (MJO), but also linked to precipitation and temperature extremes worldwide, influenced by a circumglobal Rossby wave-train along the subtropical westerly jet. Despite a close phase-lock relationship between the MJO and subseasonal EAWM precipitation, our findings indicate that the MJO itself may only play a minor role in the subseasonal EAWM variability. Given its significant impact on the S2S variability of global weather extremes, we call for coordinated community efforts to enhance the understanding and prediction of the circumglobal Rossby wave-train.

Atmospheric science↗

Meta2DB: Curated Shotgun Metagenomic Feature Sets and Metadata for Health State Prediction

Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health.

Kok, C [Lawrence Livermore National Laboratory (LL↗

A Universal Augmentation Framework for Long-Range Electrostatics in Machine Learning Interatomic Potentials

Most current machine learning interatomic potentials (MLIPs) rely on short-range approximations, without explicit treatment of long-range electrostatics. To address this, we recently developed the Latent Ewald Summation (LES) method, which infers electrostatic interactions, polarization, and Born effective charges (BECs), just by learning from energy and force training data. Here, in this study, we present LES as a standalone library, compatible with any short-range MLIP, and demonstrate its integration with methods such as MACE, NequIP, Allegro, CACE, CHGNet, and UMA. We benchmark LES-enhanced models on distinct systems, including bulk water, polar dipeptides, and gold dimer adsorption on defective substrates, and show that LES not only captures correct electrostatics but also improves accuracy. Additionally, we scale LES to large and chemically diverse data by training MACELES-OFF on the SPICE set containing molecules and clusters, making a universal MLIP with electrostatics for organic systems, including biomolecules. MACELES-OFF is more accurate than its short-range counterpart (MACE-OFF) trained on the same data set, predicts dipoles and BECs reliably, and has better descriptions of bulk liquids. By enabling efficient long-range electrostatics without directly training on electrical properties, LES paves the way for electrostatic foundation MLIPs.

Kim, Dongjin [University of California, Berkeley, ↗

Hyperdimensional computing for image classification (HDC) v1.0

This is an implementation of the hyperdimensional computing technique to classify images. It consists of a python script that trains the system for a set of images from a set of images (dataset) specified by the user. This training produces hardware configuration parameters and description vectors that are then loaded into the hardware description part of the project. The hardware description consists of hardware described in Verilog (a well known language for this purpose) that is synthesizable and can be implemented in a real chip. This hardware received the training information generated by python, and then is able to accept images to produce answers for each image on which category (class) from the pre-=trained ones the image belongs to. The hardware and python training scripts are configurable and documented. The advantage of hyperdimensional computing is its robustness to errors and the easy capability for online learning (refining the training during inference slowly over time), which this implementation supports.

Michelogiannakis, Georgios [Lawrence Berkeley Nati↗

Self-supervised and multi-fidelity learning for extended predictive soil spectroscopy

Infrared spectroscopy is a cost-effective, non-destructive, and environmentally benign technology that is increasingly recognized as an important solution for meeting the global demand for soil data. While both near-infrared (NIR) and mid-infrared (MIR) diffuse reflectance spectroscopy enable rapid estimation of soil properties, they present a significant trade-off: NIR offers superior scalability and lower operational costs, whereas MIR provides higher analytical fidelity by capturing fundamental molecular vibrations. In this study, we propose a self-supervised, multi-fidelity learning framework designed to bridge this gap. Our approach leverages large-scale MIR spectral libraries to learn a compact, transferable latent representation, into which NIR spectra are subsequently aligned for downstream prediction. The workflow consists of pretraining a latent model on a large MIR library, adapting the representation using a smaller paired NIR–MIR dataset, and evaluating generalization on an independent external test set. Across a range of chemical and physical soil properties, we found that MIR-derived embeddings improved prediction accuracy relative to baseline models that used raw MIR inputs. Predictions derived from the spectrum conversion (NIR to MIR) task did not match the performance of the original MIR spectra but were similar or superior to predictive performance of NIR-only models, suggesting the unified spectral latent space can effectively leverage the larger and more diverse MIR dataset for prediction of soil properties not well represented in current NIR libraries.

54 ENVIRONMENTAL SCIENCES↗

Direct interpolative construction of the discrete Fourier transform as a matrix product operator

The quantum Fourier transform (QFT), which can be viewed as a reindexing of the discrete Fourier transform (DFT), has been shown to be compressible as a low-rank matrix product operator (MPO) or quantized tensor train (QTT) operator. However, the original proof of this fact does not furnish a construction of the MPO with a guaranteed error bound. Meanwhile, the existing practical construction of this MPO, based on the compression of a quantum circuit, is not as efficient as possible. We present a simple closed-form construction of the QFT MPO using the interpolative decomposition, with guaranteed near-optimal compression error for a given rank. This construction can speed up the application of the QFT and the DFT, respectively, in quantum circuit simulations and QTT applications. We also connect our interpolative construction to the approximate quantum Fourier transform (AQFT) by demonstrating that the AQFT can be viewed as an MPO constructed using a different interpolation scheme.

97 MATHEMATICS AND COMPUTING↗

Ca X ML: Chemistry‐informed machine learning explains mutual changes between protein conformations and calcium ions in calcium‐binding proteins using structural and topological features

Proteins' flexibility is a feature in communicating changes in cell signaling instigated by binding with secondary messengers, such as calcium ions, associated with the coordination of muscle contraction, neurotransmitter release, and gene expression. When binding with the disordered parts of a protein, calcium ions must balance their charge states with the shape of calcium-binding proteins and their versatile pool of partners depending on the circumstances they transmit. Accurately determining the ionic charges of those ions is essential for understanding their role in such processes. However, it is unclear whether the limited experimental data available can be effectively used to train models to accurately predict the charges of calcium-binding protein variants. Here, we developed a chemistry-informed, machine-learning algorithm that implements a game theoretic approach to explain the output of a machine-learning model without the prerequisite of an excessively large database for high-performance prediction of atomic charges. We used the ab initio electronic structure data representing calcium ions and the structures of the disordered segments of calcium-binding peptides with surrounding water molecules to train several explainable models. Network theory was used to extract the topological features of atomic interactions in the structurally complex data dictated by the coordination chemistry of a calcium ion, a potent indicator of its charge state in protein. Our design created a computational tool of Ca X ML, which provided a framework of explainable machine learning model to annotate ionic charges of calcium ions in calcium-binding proteins in response to the chemical changes in an environment. Our framework will provide new insights into protein design for engineering functionality based on the limited size of scientific data in a genome space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Spectrotemporal Shaping of Attosecond X-Ray Pulses with a Fresh-Slice Free-Electron Laser

We propose a scheme allowing coherent shaping, i.e., controlling both the amplitude and phase, of attosecond x-ray pulses at free-electron lasers. Here, we show that by seeding an FEL with a short coherent seed that overfills the amplification bandwidth, one can shape the Wigner function of the pulse by controlling the undulator taper profile. The examples of controllable pulse pairs and trains, as well as isolated spectrotemporally shaped pulses with very broad bandwidths are examined in detail. Existing attosecond XFELs can achieve these experimental conditions in a two-stage cascade, in which the seed is generated by a short current spike in an electron bunch and shaped in an unspoiled region within the same bunch. We experimentally demonstrate the production and control of phase-stable pulse trains using this method at the Linac Coherent Light Source II.

47 OTHER INSTRUMENTATION↗

Reducing Frequency Bias of Fourier Neural Operators in 3D Seismic Wavefield Simulations Through Multistage Training

The recent development of neural operator (NeurOp) learning for solutions to the elastic wave equation shows promising results and provides the basis for fast large-scale simulations for different seismological applications. In this article, we use the Fourier neural operator (FNO) model to directly solve the 3D Helmholtz wave equation for fast seismic ground-motion simulations on different frequencies and show the frequency bias of the FNO model, that is, it learns the lower frequencies better comparing to the higher frequencies. To reduce the frequency bias, we adopt the multistage FNO training, that is, after training a stage 1 FNO model for estimating the ground motion, we use a second FNO model as the stage 2 to learn from the residual, which greatly reduced the errors on the higher frequencies. By adopting this multistage training, the FNO models show reduced biases on higher frequencies, which enhanced the overall results of the ground-motion simulations. Thus the multistage training FNO improves the accuracy and realism of the ground-motion simulations.

earthquakes↗

Microbiome dynamics in the congregate environment of U.S. Army Infantry training

Within military training and operational environments, individuals from diverse backgrounds share common spaces, follow structured routines and diets, and engage in physically demanding tasks. While there has been interest in leveraging microbiome features to predict and improve military health and performance, the longitudinal convergence of microbiomes in such constrained environments has not been established. To assess the degree of microbiome convergence, we performed shotgun metagenomic sequencing on swab samples from a military trainee cohort. Samples were taken across four different body sites, three timepoints, and two spatially distinct platoons. We observed evidence of convergence in one platoon, whereby similarity in microbiome composition increased over time, with numerous differentially abundant species. We found no indication of strain transfer between individuals, suggesting that convergence was influenced by external environmental factors, diet, and lifestyle. Microbial shifts observed in the convergence process included a decrease in fungal species, such as Malassezia restricta in nasal cavities, and a decrease in Prevotella species at inguinal regions across time. Shifts in multiple Corynebacterium species were also observed with varying magnitudes depending on the body site. Overall, we provide preliminary evidence of convergence of host microbial communities in military-associated environments that were distinguishable using shotgun metagenomic sequencing approaches. The data presented here on microbiome convergence, dynamics, and stability may inform risk-based mitigation in congregate military settings facilitating development of targeted microbial, dietary, or other interventions to optimize health and performance of military populations.

Biological and medical sciences↗

A Decadal Hybrid GCM Simulation Using Deep‐Learning‐Based Cloud and Convection Parameterization Generalized to a Warm Climate

A critical challenge for machine‐learning (ML) parameterization in global climate models (GCMs) is to achieve stable, accurate simulations under climates not seen during training. Previous studies have demonstrated promising offline performance and year‐long online stability in aquaplanet simulations but have encountered difficulties in real geography and under climate warming. Here we report that a GCM with real geography configuration using neural‐network‐based cloud and convection parameterization, trained exclusively with present‐day climate data, successfully performs a stable, decade‐long simulation of a warm climate with +4 K sea surface temperature (SST). The neural network (NN) is based on Han et al. (2023, https://doi.org/10.1029/2022ms003508 ) with additional inputs. The simulation captures the global precipitation distribution, surface temperatures, vertical atmospheric structures, and extreme precipitation very well, closely matching simulations from both the superparameterized CAM (SPCAM) and the conventional CAM5 in the warm climate without accuracy degradation compared to those in the baseline climate. Moreover, it produces a climate response to +4 K SST in atmospheric thermodynamic states and circulations similar to those from SPCAM and CAM5. Prognostic ablation tests on NN input variables show that the NN without convective memory as input suffers from numerical instability, and the NN without considering radiative variables and land fraction as input, or with reduced training samples produce less accurate results. To our knowledge, this is the first time an ML parameterization successfully achieves online extrapolation to a warm climate without using additional warm‐climate data for training. It demonstrates the potential of ML‐driven parameterizations for credible long‐term climate projections.

Atmosphere model↗

Multiscale Neural Networks for Approximating Green’s Functions

Neural networks (NNs) have been widely used to solve partial differential equations (PDEs) in the applications of physics, biology, and engineering. One effective approach for solving PDEs with a fixed differential operator is learning Green’s functions. However, Green’s functions are notoriously difficult to learn due to their poor regularity, which typically requires larger NNs and longer training times. In this work, we address these challenges by leveraging multiscale NNs to learn Green’s functions. Through theoretical analysis using multiscale Barron space methods and experimental validation, we show that the multiscale approach significantly reduces the necessary NN size and accelerates training.

97 MATHEMATICS AND COMPUTING↗

Towards Exascale Astrophysics of Mergers and Supernovae (TEAMS)

The TEAMS project brought together cutting-edge simulations, theoretical insights, and collaborative efforts to deepen our understanding of some of the universe’s most extreme phenomena—supernovae, neutron star mergers, and the powerful signals they emit. Using one of the largest suites of 3D supernova simulations ever conducted, researchers uncovered new insights into how massive stars explode, how those explosions vary by stellar mass, and what conditions lead to the birth of neutron stars or black holes. They also studied the radiation and gravitational wave signals emitted during these events, revealing how future observations can be used to uncover what happens deep inside collapsing stars. The team developed improved tools for modeling how light and neutrinos behave in such explosive environments, enabling more accurate predictions of what astronomers might observe. Work also explored how the chemical composition and geometry of kilonovae—the visible explosions that follow neutron star mergers—influence their signals and can reveal the origins of heavy elements like gold. These efforts not only advanced scientific knowledge, but also trained a new generation of researchers at the intersection of astrophysics, computational science, and nuclear theory.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

15-minute Parker River gap-filled tide height and salinity data, PIE LTER, Plum Island Sound, MA (2014–2023), for ELM PFLOTRAN modeling

This dataset contains 15-minute tide height and salinity data from the Typha site along the Parker River, part of the Plum Island Ecosystems Long Term Ecological Research (PIE LTER) site in Plum Island Sound, Massachusetts (MA) 2014-2023. Tide height (in NAVD88) was compiled from measurements conducted at the mouth of Plum Island Sound and corrected for time lags. Gap-filling of missing periods were done by fitting tidal constituents to the time series. Salinity was measured (and is stored on ESS DIVE ) in 2022 and 2023 using HOBO U24-002 conductivity loggers. River discharge is the most important control on tidal river water salinity at the location (Vallino & Hopkinson, 1998). An artificial neural network was trained to predict river water salinity at the location using Parker River discharge (USGS station 01101000, Parker River at Byfield, MA) and gap-filled salinity observations from a long-term monitoring station ca. 3km downstream from the Typha site (LTER station ‘Middle Road’) as input variables to create continuous time series information. The data set was used in the spin up and simulations of a land surface model coupled to a biogeochemical reaction network (ELM PFLOTRAN) assessing impacts of hydrology and salinity input on methane fluxes in 2022 and 2023 (Sulman et al., 2024). Metadata files ELMPFLOTRAN_tide_salinity_dd.csv and ELMPFLOTRAN_tide_salinity_flmd.csv provide details on site location, data variables, and QA/QC methods .

54 ENVIRONMENTAL SCIENCES↗

Discriminative versus generative approaches to simulation-based inference

Most of the fundamental, emergent, and phenomenological parameters of particle and nuclear physics are determined through parametric template fits. Simulations are used to populate histograms which are then matched to data. This approach is inherently lossy, since histograms are binned and low-dimensional. Deep learning has enabled unbinned and high-dimensional parameter estimation through neural likelihood(-ratio) estimation. We compare two approaches for neural simulation-based inference (NSBI): one based on discriminative learning (classification) and one based on generative modeling. These two approaches are directly evaluated on the same datasets, with a similar level of hyperparameter optimization in both cases. In addition to a Gaussian dataset, we study NSBI using a Higgs boson dataset from the FAIR Universe Challenge. We find that both the direct likelihood and likelihood ratio estimation are able to effectively extract parameters with reasonable uncertainties. For the numerical examples and within the set of hyperparameters studied, we found that the likelihood ratio method is more accurate and/or precise. Both methods have a significant spread from the network training and would require ensembling or other mitigation strategies in practice.

high energy physics↗

tsSLOPE

This project implements the interface that reads a machine learning (ML) model trained in Python to be used in Julia to inform a JuMP optimization model.

Chiang, Nai Yun [Lawrence Livermore National Labor↗

Shaping the FutureWorkforce: Challenges and Lessons Learned in HPC Education from National Labs and Computing Centers

Workforce training at national laboratories and computing centers is essential and typically falls into two categories: foundational training for newcomers and advanced training for experienced users. Foundational topics—such as version control, build systems, and basic HPC usage—are largely transferable across institutions, while cluster-specific training varies due to differences in hardware, job schedulers, and local workflows. Training on emerging technologies is split between hardware-specific content and broadly applicable programming paradigms. Here, to reduce redundancy and increase impact, national labs, computing centers, and vendors are collaborating through initiatives like the HPC Training Working Group to share best practices, co-develop materials, and broaden outreach. These coordinated efforts aim to make HPC training more accessible, scalable, and consistent across the community.

HPC↗