Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model data volumes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Tidal Currents in San Juan Archipelago, Washington

Re-analyzed acoustic Doppler current profiler (ADCP) data originally collected by NOAA CO-OPS (Center for Operational Oceanographic Products and Services) and equivalent point data from Pacific Northwest National Laboratory's FVCOM (Finite Volume Community Ocean Model) model of the region. Data are processed to products describing characteristics of tidal currents relevant to tidal turbines, as well as power output estimates for a notional turbine deployed from a surface platform or from the seabed at each location. These data underpin the results presented in their associated paper - see below.

16 TIDAL AND WAVE POWER↗

BISON High Burnup Structure Modeling Capabilities Validated with a Selection of the Halden IFA-650 Rods

The U.S. Department of Energy (DOE)’s Nuclear Energy Advanced Modeling and Simulation (NEAMS)program aims to develop predictive capabilities using computational methods for the analysis and design of advanced reactor and fuel cycle systems. This program has been supporting the development of BISON, a high-fidelity and high-resolution fuel performance tool at the engineering scale. This document continues analysis and refinement of capabilities added to BISON early this calendar year in regards to the incorporation of capabilities applicable to extended burnups in response to industry interest. Details are provided on high burnup thermal conductivity models, a refitting of the high-burnup structure (HBS) porosity formation model to include additional data, the coupling of the HBS volume fraction model to thermal conductivity and fine fragmentation models, and validation activities. The IFA-650.4 and IFA-650.9 loss of coolant accident (LOCA) analyses are revisited with the latest developments in this report. A new validation case, IFA-650.14 has also been added to the BISON test suite.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

FedCSpc: A Cross-Silo Federated Learning System With Error-Bounded Lossy Parameter Compression

Cross-Silo federated learning is widely used for scaling deep neural network (DNN) training over data silos from different locations worldwide while guaranteeing data privacy. Communication has been identified as the main bottleneck when training large-scale models due to large-volume model parameters and gradient transmission across public networks with limited bandwidth. Most previous works focus on gradient compression, while limited work tries to compress parameters that can not be ignored and extremely affect communication performance during the training. Here, to bridge this gap, we propose FedCSpc: an efficient cross-silo federated learning system with an XAI-driven adaptive parameter compression strategy for large-scale model training. Our work substantially differs from existing gradient compression techniques due to the distinct data features of gradient and parameter. The key contributions of this paper are fourfold. (1) Our designed FedCSpc proposes to compress the parameter during the training using the state-of-the-art error-bounded lossy compressor – SZ3. (2) We develop an adaptive compression error bound adjustment algorithm to guarantee the model accuracy effectively. (3) We exploit an efficient approach to utilize the idle CPU resources of clients to compress the parameters. (4) We perform a comprehensive evaluation with a wide range of models and benchmarks on a GPU cluster with 65 GPUs. Results show that FedCSpc can achieve the same model accuracy as FedAvg while reducing the data volume of parameters and gradients in communication by up to 7.39× and 288×, respectively. With 32 clients on a 4 Gb size model, FedCSpc significantly outperforms FedAvg in wall-clock time in the emulated WAN environment (at the bandwidth of 1 Gbps or lower without loss of generality).

SZ3↗

Multitask graph neural networks for elastoplastic response prediction in dual-phase polycrystals

Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.

Crystal plasticity↗

A Spectacular Viscoelastic Model Calibration for 828/D230/Alox Generated from Legacy Sandia Data

The nonlinear viscoelastic Spectacular model is calibrated to the thermo-mechanical behavior of 828/D230/Alox with an alox volume fraction of 20 %. Legacy experimental data from Sandia’s polymer properties database (PPD) is used to calibrate the model. Based on known densities of the epoxy 828/D230 and the alox filler, the alox volume fractions listed on the PPD were likely reported incorrectly. The alox volume fractions are recalculated here. Using the recalculated alox volume fractions, the PPD contains experimental data for 828/D230/Alox with alox volume fractions of 16 %, 24 %, and 33 %, so the thermo-mechanical behavior at 20 % alox volume fraction is estimated by interpolating between the bounding cases of of 16 % and 24 %. Because the Spectacularmodel can be fairly challenging to calibrate, the calibration procedure is described in detail. Several of the calibration steps involve inverse parameter identification, where an experiment is simulated and parameters are iteratively updated until the model response matches the experimental data. As the PPD does not fully describe all experimental procedures, the experimental simulations use assumed thermal and mechanical loading rates that are typical for the viscoelastic characterization of epoxies. Spectacular uses four independent relaxation functions related to volumetric (ƒ 1 ), shear (ƒ 2 ), thermal strain (ƒ 3 ), and thermal relaxations (ƒ 4 ). The previous SPEC model form, also known as the universal_polymer model, uses two independent relaxation functions related to volumetric and thermal relaxation (ƒ ν = ƒ 1 = ƒ 3 = ƒ 4 ) and shear relaxation (ƒ s = ƒ 2 ). The two constitutive choices are briefly evaluated here, where it is found that the four relaxation function approach of Spectacular was better suited for fitting the coefficient of thermal expansion during both heating and cooling.

36 MATERIALS SCIENCE↗

Two-Stream Multi-Channel Convolutional Neural Network for Multi-Lane Traffic Speed Prediction Considering Traffic Volume Impact

Traffic speed prediction is a critically important component of intelligent transportation systems. Recently, with the rapid development of deep learning and transportation data science, a growing body of new traffic speed prediction models have been designed that achieved high accuracy and large-scale prediction. However, existing studies have two major limitations. First, they predict aggregated traffic speed rather than lane-level traffic speed; second, most studies ignore the impact of other traffic flow parameters in speed prediction. To address these issues, the authors propose a two-stream multi-channel convolutional neural network (TM-CNN) model for multi-lane traffic speed prediction considering traffic volume impact. In this model, the authors first introduce a new data conversion method that converts raw traffic speed data and volume data into spatial–temporal multi-channel matrices. Then the authors carefully design a two-stream deep neural network to effectively learn the features and correlations between individual lanes, in the spatial–temporal dimensions, and between speed and volume. Accordingly, a new loss function that considers the volume impact in speed prediction is developed. A case study using 1-year data validates the TM-CNN model and demonstrates its superiority. This paper contributes to two research areas: (1) traffic speed prediction, and (2) multi-lane traffic flow study.

97 MATHEMATICS AND COMPUTING↗

Robustness of deep learning algorithms in astronomy -- galaxy morphology studies

Deep learning models are being increasingly adopted in wide array of scientific domains, especially to handle high-dimensionality and volume of the scientific data. However, these models tend to be brittle due to their complexity and overparametrization, especially to the inadvertent adversarial perturbations that can appear due to common image processing such as compression or blurring that are often seen with real scientific data. It is crucial to understand this brittleness and develop models robust to these adversarial perturbations. To this end, we study the effect of observational noise from the exposure time, as well as the worst case scenario of a one-pixel attack as a proxy for compression or telescope errors on performance of ResNet18 trained to distinguish between galaxies of different morphologies in LSST mock data. We also explore how domain adaptation techniques can help improve model robustness in case of this type of naturally occurring attacks and help scientists build more trustworthy and stable models.

79 ASTRONOMY AND ASTROPHYSICS↗

Prioritizing the selection of CMIP6 model ensemble members for downscaling projections of CONUS temperature and precipitation

Given the mismatch between the large volume of data archived for the sixth phase of the Coupled Model Intercomparison Project (CMIP6) and limited personnel and computational resources for downscaling, only a small fraction of the CMIP6 archive can be downscaled. In this work, we develop an approach to robustly sample projected hydroclimate states in CMIP6 for downscaling to test whether the selection of a single initial condition (IC) ensemble member from each CMIP6 model is sufficient to span the range of modeled hydroclimate over the conterminous United States (CONUS) and CONUS sub-regions. We calculate the pattern-centered root mean square difference of IC ensemble member anomalies relative to each model’s historical climatology for shared socioeconomic pathway (SSP) projections over 30-year time periods and compare the ratio of inter-model to intra-model variability for this metric. Regardless of SSP, inter-model variability is generally much greater than intra-model variability at the scales of the CONUS as a whole, as well as for most CONUS sub-regions. However, for some variables and scenarios, inter- and intra-model variability are similar at sub-CONUS scales, indicating that selecting a single IC ensemble member per model may be sufficient to sample the range of projected hydroclimate states in the 21st Century across CONUS, but for specific regions and variables, more careful selection of ensemble members may be necessary. Regionally resolved Taylor diagrams identify where more IC ensemble member downscaling efforts should be focused if resources are available to do so. Our results suggest that, with parsimonious sampling, the requisite computational expense of downscaling temperature and precipitation fields over the CONUS for subsequent CMIP activities may increase only marginally despite the great increase in data volumes with each successive CMIP phase.

54 ENVIRONMENTAL SCIENCES↗

A Science Gateway for the Repeatable Analysis of Machine Learning Predicted Gravity Anomalies

In recent years, deep learning has become an increasingly popular alternative for modeling in geoscience applications due to its scalability and efficiency. However, the interpretability, compute, data volume, and hyperparameter tuning requirements of deep learning models make development and monitoring difficult. Furthermore, model explainability and communicating results obtained by these models to users or domain experts is a challenge, as domain experts in geoscience also need to have a deep understanding of how those models function in order to support their scientific works. Here, we describe a science gateway and machine learning pipeline for predicting gravity anomalies from geophysical data. The gateway, built on open-source technologies, provides a holistic view of the pipeline through interactive visualizations aimed at enabling efficient exploratory data analysis. The repeatability, reproducibility, and monitoring capabilities of this overall system allow us to iterate and analyze at scale. Using this pipeline and gateway, we can repeatedly produce accurate high-resolution gravity anomaly datasets. By describing the underlying technologies, implementation, and results, here we provide a foundation for the broader adoption of science gateways into cross-cutting geoscience and machine learning research projects as a means to improve the scientific discovery and collaboration in the geophysics and computational sciences community.

58 GEOSCIENCES↗

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

MFA-DVR: direct volume rendering of MFA models

3D volume rendering is widely used to reveal insightful intrinsic patterns of volumetric datasets across many domains. However, the complex structures and varying scales of volumetric data can make efficiently generating high-quality volume rendering results a challenging task. Multivariate functional approximation (MFA) is a new data model that addresses some of the critical challenges: high-order evaluation of both value and derivative anywhere in the spatial domain, compact representation for large-scale volumetric data, and uniform representation of both structured and unstructured data. Here, in this paper, we present MFA-DVR, the first direct volume rendering pipeline utilizing the MFA model, for both structured and unstructured volumetric datasets. We demonstrate improved rendering quality using MFA-DVR on both synthetic and real datasets through a comparative study. We show that MFA-DVR not only generates more faithful volume rendering than using local filters but also performs faster on high-order interpolations on structured and unstructured datasets. MFA-DVR is implemented in the existing volume rendering pipeline of the Visualization Toolkit (VTK) to be accessible by the scientific visualization community.

97 MATHEMATICS AND COMPUTING↗

Computer-aided design file format for additive manufacturing and methods of file generation

A method of generating a tessellated output file comprising receiving a computer-aided design (CAD) model file defining a CAD model including a plurality of vertices, a plurality of curves, a plurality of surfaces, and at least one volume; generating base polygon data defining the CAD model for additive manufacturing, the base polygon data including a plurality of connected polygons, each polygon including a plurality of nodes, a plurality of edges, and a face; generating data for a vertex metadata container including a listing of CAD model vertices and one polygon node associated with each CAD model vertex; generating data for a curve metadata container including a listing of CAD model curves and at least one polygon edge associated with each CAD model curve; and generating data for a surface metadata container including a listing of CAD model surfaces and at least one polygon face associated with each CAD model surface.

Vernon, Gregory John↗

Machine learning assisted rediscovery of methane storage and separation in porous carbon from material literature

Porous carbon (PC) has been widely regarded as one of the most promising absorbents for methane storage. Studies show that its uptake capacity and selectivity highly depend on textural structures. Although much effort has been made, unveiling their detailed structure-performance relationship remains a challenge. Here, we propose an innovative study where, with the assistance of machine learning, the hidden relationship of the textural structures of PC with the methane uptake and separation can be derived from existing data in material literature. Machine learning models were trained by the data, including specific surface area, micropore volume, mesopore volume, temperature, and pressure as the input variables and methane uptake as the output variable for prediction. Among the tested models, the multilayer perceptron (MLP) shows the highest accuracy in predicting the methane uptake. In addition, the model enables to automatically construct a uptake performance map in terms of micropore volume and mesopore volume. The obtained MLP model was also extended to explore the CO 2 /CH 4 selectivity by retraining it with the data collected from literature of PC for the CO 2 uptake. Finally, the constructed 2D selectivity map shows that the high selectivity can be achieved in the low CH 4 uptake region.

42 ENGINEERING↗

A Perspective on Data and Privacy for AI in Healthcare [Industrial and Governmental Activities]

As large language models continue to push the bounds of AI model size, they are also being trained on unprecedented volumes of data. While individual hospitals are estimated to produce petabytes of data per year, only a small fraction is currently being used for developing AI models. Additionally, with such data resources available, healthcare is well-positioned to benefit from the current trends in AI. Moreover, the inherently multi-modal and longitudinal nature of clinical data – from omics to imaging to unstructured notes – provides a fertile ground for the development and application of cutting-edge architectures like foundation models.

Gounley, John [Oak Ridge National Laboratory (ORNL↗

Final technical report for DE-SC0022255: Discovering Physically Meaningful Structures from Climate Extreme Data

The past two decades have witnessed natural disasters and extreme weather events that affect millions of people. At the same time, the data volume from high-resolution climate models, satellite, in-situ and ground-based measurements have substantially increased to petabyte scales. These new and readily accessible datasets create the previously missing pipeline required for scientific machine learning (ML) and therefore new opportunities for improved understanding and prediction capability of climate extreme events. This project developed a deep latent variable model framework to discover physically meaningful hidden structures from high-dimensional, spatiotemporal climate extreme data.

97 MATHEMATICS AND COMPUTING↗

DDStore: Distributed Data Store for Scalable Training of Graph Neural Networks on Large Atomistic Modeling Datasets

Graph neural networks (GNNs) are a class of Deep Learning models used in designing atomistic materials for effective screening of large chemical spaces. To ensure robust prediction, GNN models must be trained on large volumes of atomistic data on leadership class supercomputers. Even with the advent of modern architectures that consist of multiple storage layers that include node-local NVMe devices in addition to device memory for caching large datasets, extreme-scale model training faces I/O challenges at scale.We present DDStore, an in-memory distributed data store designed for GNN training on large-scale graph data. DDStore provides a hierarchical, distributed, data caching technique that combines data chunking, replication, low-latency random access, and high throughput communication. DDStore achieves near-linear scaling for training a GNN model using up to 1000 GPUs on the Summit and Perlmutter supercomputers, and reaches up to a 6.15x reduction in GNN training time compared to state-of-the-art methodologies.

Choi, Jong Youl↗

Data-driven Whitney forms for structure-preserving control volume analysis

Control volume analysis models physics via the exchange of generalized fluxes between subdomains. Here, we introduce a scientific machine learning framework adopting a partition of unity architecture to identify physically-relevant control volumes, with generalized fluxes between subdomains encoded via Whitney forms. The approach provides a differentiable parameterization of geometry which may be trained in an end-to-end fashion to extract reduced models from full field data while exactly preserving physics. The architecture admits a data-driven finite element exterior calculus allowing discovery of mixed finite element spaces with closed form quadrature rules. An equivalence between Whitney forms and graph networks reveals that the geometric problem of control volume learning is equivalent to an unsupervised graph discovery problem. The framework is developed for manifolds in arbitrary dimension, with examples provided for H(div) problems in $\mathbb{R}$ establishing convergence and structure preservation properties. Finally, we consider a lithium-ion battery problem where we discover a reduced finite element space encoding transport pathways from high-fidelity microstructure resolved simulations. The approach reduces the 5.89M finite element simulation to 136 elements while reproducing pressure to under 0.1% error and preserving conservation.

97 MATHEMATICS AND COMPUTING↗

Enhancing the accuracy and generality of the Debye–Grüneisen Model: Optimizing the volume dependence for accurate predictions across varied compositions

In this work, we have introduced an optimized Debye-Grüneisen model that revolutionizes the determination of the Debye temperature and Grüneisen parameters. Unlike conventional methods, our model requires only the 0 K energy volume data for a material as input, eliminating the need to determine the bulk modulus and its pressure derivative, which often pose challenges due to numerical uncertainties. This unique feature sets our model apart from existing approaches and streamlines the process, enabling accurate predictions of thermal expansion behavior across various materials. To demonstrate its effectiveness, we showcase its excellent agreement with measured coefficients of thermal expansion (CTE) for the nickel-cobalt-chromium-aluminum-yttrium (Ni-Co-Cr-Al-Y) bond-coating system. Additionally, we apply our approach by conducting a high-throughput search for potential bond-coating materials among 90,000 compositions within the aluminum-cobalt-chromium-iron-nickel (Al-Co-Cr-Fe-Ni) system. From this extensive search, four compositions are synthesized, and the measured CTE values agree very well with theoretical predictions, hence validating our approach. In conclusion, the current optimized Debye-Grüneisen model combined with Density Functional Theory (DFT)-based thermodynamic database enables reliable and efficient high-throughput calculations of CTE of of a material without expensive phonon calculations.

Bond coating materials↗