Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “load data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Block encoding of the three-dimensional heterogeneous Poisson equation with application to fracture flow

Quantum linear system (QLS) algorithms offer the potential to solve large-scale linear systems exponentially faster than classical methods. However, applying QLS algorithms to real-world problems remains challenging due to issues such as state preparation, data loading, and efficient information extraction. In this work, we study the feasibility of applying QLS algorithms to solve discretized three-dimensional (3D) heterogeneous Poisson equations, with specific examples relating to groundwater flow through geologic fracture networks. We explicitly construct a block encoding for the 3D heterogeneous Poisson matrix by leveraging the sparse local structure of the discretized operator. While classical solvers benefit from preconditioning, we show that block encoding the system matrix and preconditioner separately does not improve the effective condition number that dominates the QLS run-time. This differs from classical approaches where the preconditioner and the system matrix can often be implemented independently. Nevertheless, due to the structure of the problem in three dimensions, the quantum algorithm achieves a run-time of 𝑂⁡(𝑁 2/3 polylog 𝑁 ⋅log (1/𝜖)), outperforming the best classical methods (with run times of 𝑂⁡(𝑁⁢log 𝑁 ⋅log (1/𝜖))) and offering exponential memory savings. These results highlight both the promise and limitations of QLS algorithms for practical scientific computing, and point to effective condition-number reduction as a key barrier in achieving quantum advantages.

58 GEOSCIENCES↗

Performance Profile of Transformer Fine-Tuning in Multi-GPU Cloud Environments

The study presented here focuses on performance characteristics and trade-offs associated with running machine-learning tasks in multi-GPU environments on both on-site cloud computing resources and commercial cloud services (Azure). Specifically, this study examines these tradeoffs by examining the performance of training and fine-tuning of transformer-based deep-learning (DL) networks on clinical notes and data, a task of critical importance in the medical domain. To this end, we perform DL-related experiments on the widely deployed NVIDIA V100 GPUs and on the newer A100 GPUs connected via NVLink or PCIe. This study analyzes the execution time of major operations to train DL models and investigate popular options to optimize each of them. We examine and present the findings on the impacts that various operations (e.g. data loading into GPUs, training, fine-tuning), optimizations, and system configurations (single vs. multi-GPU, NVLink vs. PCIe) have on the overall training performance.

Begoli, Edmon↗

Using Artificial Intelligence to Improve Reliability and Operational Efficiency of Small-Scale Hydroelectric Distributed Generation

Reliability and resilience are critical concerns for distributed generation (DG) at the rural electric level. The integration of renewable energy sources, such as small-scale hydroelectric distributed generators (hydro DGs), introduces operational challenges, particularly regarding aging infrastructure and grid stability. Artificial Intelligence (AI)-driven Machine Learning (ML) models and applications of Large Language Models (LLMs) offer promising solutions for optimizing DG operations and enhancing resilience. This paper explores AI-based models for improving efficiency, fault resolution, and outage mitigation in small-scale hydro DGs. Furthermore, it highlights the development of a centralized, AI-powered information portal for rural electric cooperatives and municipalities. The research evaluates hydro DG plant models and discusses the applicability of AI-powered question-answering tools for real-time operations, focusing on statistical data, load flow, voltage regulation, and generation power. The findings demonstrate AI’s potential to transform DG management to ensure greater stability and resilience in rural electric grids.

Bhattacharyya, Arjun [ORNL] (ORCID:000900060976046↗

An Adaptive-Importance-Sampling-Enhanced Bayesian Approach for Topology Estimation in an Unbalanced Power Distribution System

The reliable operation of a power distribution system relies on a good prior knowledge of its topology and its system state. Although crucial, due to the lack of direct monitoring devices on the switch statuses, the topology information is often unavailable or outdated for the distribution system operators for real-time applications. Apart from the limited observability of the power distribution system, other challenges are the nonlinearity of the model, the complicated, unbalanced structure of the distribution system, and the scale of the system. To overcome the above challenges, we, in this paper, propose a Bayesian-inference framework that allows us to simultaneously estimate the topology and the state of a three-phase, unbalanced power distribution system. Specifically, by using the very limited number of measurements available that are associated with the forecast load data, we efficiently recover the full Bayesian posterior distributions of the system topology under both normal and outage operation conditions. This is performed through an adaptive importance sampling procedure that greatly alleviates the computational burden of the traditional Monte-Carlo (MC)-sampling-based approach while maintaining a good estimation accuracy. The simulations conducted on the IEEE 123-bus test system and an unbalanced 1282-bus system reveal the excellent performances of the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

TEE-ACM2

TEE-ACM2 is a library that computes matrix chain multiplication efficiently on GPUs, using a blocking strategy to load data. The library will provide an energy efficient algorithm for chain Matrix Multiplication, by minimizing both computations and off chip data transfers on the GPUs.

Lim, Hyun↗

HydraGNN 2.0

HydraGNN is an Oak Ridge National Laboratory (ORNL)-branded implementation of distributed multi-tasking graph neural networks that supports several scientific applications within the ORNL portfolio to support the US-DOE mission. New or improved capabilities included in v2.0.0 release are as follows: 1) Enhancement in message passing layers through class inheritance 2) Adding transformation to ensure translation and rotation invariance 3) Supporting various optimizers 4) Atomic descriptors 5) Integration with continuous CI test 6) Distributed printouts and timers 7) Profiling 8) Support of ADIOS2 for scalable data loading 9) Large-scale system support, including Summit (ORNL) and Perlmutter (NERSC)

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

pnnl/slim

Open source release of Python Systems Library which contains benchmark datasets, system emulators, and data loading codes.

Tuor, Aaron↗

MalGen

MalGen is helper code that includes a script for loading data in format for PyTorch neural network models. It also includes example models. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3078 O

Johnson, NicholasT↗

ParFlow Sand Tank: A tool for groundwater exploration

The ParFlow Sand Tank model is an open source application designed to allow users to interactively simulate and visualize groundwater movement through the subsurface. The app is designed for both research and education; teaching hydrogeology concepts and making it easy explore and run sophisticated groundwater simulations. Our goal is to support increased accessibility and usability of research grade hydrology tools for research and teaching. The Sand Tank application simulates groundwater and surface water fluxes as well as contaminant transport in real time using the integrated physical hydrology model ParFlow (Kollet & Maxwell, 2006; Maxwell & Miller, 2005; Osei-Kuffuor et al., 2014) and the particle tracking code EcoSlim (Maxwell et al., 2019). ParFlow is a numerical hydrology model that simulates spatially distributed groundwater and surface water flow. It is a well established research tool with more than 90 publications documenting its development use to advance our understanding of groundwater dynamics and groundwater surface water interactions from the hillslope to the continental scale e.g. (Condon et al., 2020; Condon & Maxwell, 2019; Maxwell & Condon, 2016). It is designed for efficient parallel computation and has been run on many platforms spanning from laptops to supercomputers. However, one of the challenges of ParFlow is that it requires significant training and hydrologic expertise to develop simulations. The Sand Tank application makes this model accessible to anyone for education and exploration. Our application uses ParFlow for its simulation backend and ParaView for the data loading and processing. The communication infrastructure relies on the ParaViewWeb framework. We use model templates deployed in Docker images to setup the Sand Tank framework. Users can build the application locally or interact with it through our web deployment. When interacting with a template users can interactively change model parameters like subsurface processes or pump/inject water into the subsurface and watch the system respond to their changes in real time as the simulation runs. Additionally, our template setup will allow more advanced users to build custom templates of increasing complexity for both research and educational purposes.

54 ENVIRONMENTAL SCIENCES↗

Stabilized Hyperfoam Modeling of the General Plastics EF4003 (3 PCF) Flexible Foam

Constitutive model parameterizations for the General Plastics EF4003 low density 3 pound per cubic foot are needed for design and qualification purposes in normal and abnormal mechanical simulations. The material is expected to be deformed in two ways: first during preloading, and second under impact conditions of the system (transient dynamic). All analyses are to be performed at room temperature. The goal is to provide the analysis community a robust constitutive model parameterization to represent the compression behavior of the EF4003 foam from small deformations up to massive compressive deformations when the foam is densifying. It is worth noting the EF4003 exhibits anisotropy in its stress-strain behavior between the rise and transverse directions (See figure 2.8c-d) as well as plateau behavior that is very likely to cause material stability issues, due to the buckling transition, (and has historically done so) when using Sandia’s current workhorse models for flexible foams, Hyperfoam and Flex Foam. A Stability-informed Hyperfoam parameterization procedure is developed and executed to calibrate a hyperfoam model for the EF4003 room temperature, transversely loaded data. A rise orientation parameterization was not attempted due to localization in the experiments.

36 MATERIALS SCIENCE↗

Interpretable Net Load Forecasting Using Smooth Multiperiodic Features

We consider the problem of forecasting net load over a horizon such as one day, using a trailing window of past net load values as well as date and time. We focus on three variations on this problem: point forecasts, marginal quantile forecasts, and generating conditional samples of the future value. We propose a method that relies on linear regression using some custom engineered time-based features to capture multiple periodicities, such as daily, weekly, and seasonal, and their interactions. Our proposed models are readily interpretable, and rely on efficient and reliable convex optimization [1] to fit. We illustrate our method on four years worth of hourly net load data, comparing predictions made with various subsets of the features.

Ogut, Mehmet G↗

TEAMER – Enhanced Flow Measurement for Aquantis Tidal Turbine Test

The AQ10 is a floating, two-bladed, passive yawing tidal turbine developed by Aquantis that has a 10-meterrotor diameter, 160 kW rating, and employs reliable off-the-shelf powertrain and power conversion hardware. Aquantis is planning on-water turbine power performance and loads (blade loading and thrust)testing, where the turbine will be pushed through still water up to 4 knots and placed in a ‘station keeping’ tow in a tidal race up to its rated speed. On-water testing will be conducted using vessels and floating platforms on the sea surface to improve ease of testing and reduce disturbance to the environment. In this TEAMER project, Pacific Northwest National Laboratory (PNNL) will conduct water velocity and turbulence measurements in front of the turbine during on-water testing using acoustic Doppler instrumentation. By measuring both the tidal current flowing past the turbine and the resulting electrical power output, test results will provide a power curve (power vs flow speed) for the turbine up to rated power. Measurements of turbulence and velocity shear in front of the rotor will also provide information to assess the structural response of the rotor blades. With this analysis, Aquantis can use the performance and loads data to validate Tidal Bladed and OpenFAST simulations of the measured operating conditions. Measuring the power performance of a prototype turbine is a valuable step to improving device development and conducting a complete power performance assessment to IEC/TS 62600-200 standards in the future.

16 TIDAL AND WAVE POWER↗

Benchmarking DAOS Filesystem on Aurora

We benchmark the DAOS filesystem on Argonne's Aurora supercomputer (127 nodes, 4,064 targets) using fio, IOR, mdtest, and IO500 to characterize I/O and metadata performance across the DFS API and DFuse+POSIX. Single-client fio shows POSIX bandwidth saturating at 1–2 MiB I/O sizes, with write-heavy workloads outperforming reads. Multi-node IOR shows DFS bandwidth scaling well up to ~32 tasks/node, with write latency growing faster than read latency. An 8-node IO500 evaluation shows DFS achieving ~5x higher bandwidth and ~190x higher IOPS than POSIX. Results indicate DAOS is well-suited to read-heavy workloads like AI training data loading, given appropriately sized transfers and concurrency.

George, Rebecca [College of William and Mary, Will↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

Modeling HIV-1 Within-Host Dynamics After Passive Infusion of the Broadly Neutralizing Antibody VRC01

VRC01 is a broadly neutralizing antibody that targets the CD4 binding site of HIV-1 gp120. Passive administration of VRC01 in humans has assessed the safety and the effect on plasma viremia of this monoclonal antibody (mAb) in a phase 1 clinical trial. After VRC01 infusion, the plasma viral load in most of the participants was reduced but had particular dynamics not observed during antiretroviral therapy. In this paper, we introduce different mathematical models to explain the observed dynamics and fit them to the plasma viral load data. Based on the fitting results we argue that a model containing reversible Ab binding to virions and clearance of virus-VRC01 complexes by a two-step process that includes (1) saturable capture followed by (2) internalization/degradation by phagocytes, best explains the data. This model predicts that VRC01 may enhance the clearance of Ab-virus complexes, explaining the initial viral decay observed immediately after antibody infusion in some participants. Because Ab-virus complexes are assumed to be unable to infect cells, i.e., contain neutralized virus, the model predicts a longer-term viral decay consistent with that observed in the VRC01 treated participants. By assuming a homogeneous viral population sensitive to VRC01, the model provides good fits to all of the participant data. However, the fits are improved by assuming that there were two populations of virus, one more susceptible to antibody-mediated neutralization than the other.

60 APPLIED LIFE SCIENCES↗

Testing Protocol Development for the Fracture Toughness of Parts Built with Big Area Additive Manufacturing

The mechanical testing of additively manufactured parts has largely relied on the existing standards developed for traditional manufacturing. While this approach leverages the investment made in current standards development, it inaccurately assumes that the mechanical response of additive manufacturing (AM) parts is identical to that of parts manufactured through traditional processes. When considering thermoplastic, material extrusion AM, the differences in response can be attributed to an AM part’s inherent inhomogeneity caused by porosity, interlayer zones, and surface texture. Additionally, the interlayer bonding of parts printed with large-scale AM is difficult to adequately assess, as much testing is performed such that stress is distributed across many layer interfaces; therefore, the lack of AM-specific standards to assess interlayer bonding is a significant research gap. To quantify interlayer bonding via fracture toughness, double cantilever beam (DCB) testing has been used for some AM materials, and DCB has been generally used for a variety of materials including metal, wood, and laminates. Mode I DCB testing was performed on thermoplastic matrix composites printed with Big Area Additive Manufacturing (BAAM). Of particular interest was the notch shape and deflection speed during testing. The results examine the differences when using two notch types and three deflection speeds. The testing method introduced by the following paper differentiates itself from the ones described in the standards used by modernizing the methodology. This was conducted with the introduction of Digital Image Correlation (DIC) to gather displacement and load data simultaneously without human intervention.

Polymer Science↗