Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “input data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Machine learning of high dimensional data on a noisy quantum processor

Abstract Quantum kernel methods show promise for accelerating data analysis by efficiently learning relationships between input data points that have been encoded into an exponentially large Hilbert space. While this technique has been used successfully in small-scale experiments on synthetic datasets, the practical challenges of scaling to large circuits on noisy hardware have not been thoroughly addressed. Here, we present our findings from experimentally implementing a quantum kernel classifier on real high-dimensional data taken from the domain of cosmology using Google’s universal quantum processor, Sycamore. We construct a circuit ansatz that preserves kernel magnitudes that typically otherwise vanish due to an exponentially growing Hilbert space, and implement error mitigation specific to the task of computing quantum kernels on near-term hardware. Our experiment utilizes 17 qubits to classify uncompressed 67 dimensional data resulting in classification accuracy on a test set that is comparable to noiseless simulation.

97 MATHEMATICS AND COMPUTING↗

OpenET: Filling a Critical Data Gap in Water Management for the Western United States

The lack of consistent, accurate information on evapotranspiration (ET) and consumptive use of water by irrigated agriculture is one of the most important data gaps for water managers in the western United States (U.S.) and other arid agricultural regions globally. The ability to easily access information on ET is central to improving water budgets across the West, advancing the use of data-driven irrigation management strategies, and expanding incentive-driven conservation programs. Recent advances in remote sensing of ET have led to the development of multiple approaches for field-scale ET mapping that have been used for local and regional water resource management applications by U.S. state and federal agencies. The OpenET project is a community-driven effort that is building upon these advances to develop an operational system for generating and distributing ET data at a field scale using an ensemble of six well-established satellite-based approaches for mapping ET. Key objectives of OpenET include: Increasing access to remotely sensed ET data through a web-based data explorer and data services; supporting the use of ET data for a range of water resource management applications; and development of use cases and training resources for agricultural producers and water resource managers. Here we describe the OpenET framework, including the models used in the ensemble, the satellite, meteorological, and ancillary data inputs to the system, and the OpenET data visualization and access tools. We also summarize an extensive intercomparison and accuracy assessment conducted using ground measurements of ET from 139 flux tower sites instrumented with open path eddy covariance systems. Results calculated for 24 cropland sites from Phase I of the intercomparison and accuracy assessment demonstrate strong agreement between the satellite-driven ET models and the flux tower ET data. For the six models that have been evaluated to date (ALEXI/DisALEXI, eeMETRIC, geeSEBAL, PT-JPL, SIMS, and SSEBop) and the ensemble mean, the weighted average mean absolute error (MAE) values across all sites range from 13.6 to 21.6 mm/month at a monthly timestep, and 0.74 to 1.07 mm/day at a daily timestep. At seasonal time scales, for all but one of the models the weighted mean total ET is within ±8% of both the ensemble mean and the weighted mean total ET calculated from the flux tower data. Overall, the ensemble mean performs as well as any individual model across nearly all accuracy statistics for croplands, though some individual models may perform better for specific sites and regions. We conclude with three brief use cases to illustrate current applications and benefits of increased access to ET data, and discuss key lessons learned from the development of OpenET.

54 ENVIRONMENTAL SCIENCES↗

Visualization Tool for Comparing Low-Carbon Energy Options

This project developed a web based visualization tool for comparing current and future nuclear fuel cycle options to low-carbon and conventional energy technologies in the United States. While nuclear power is well established as the gold-standard for baseload power production and also as the technology with the lowest overall carbon intensity of any commercial form of electricity generation. This project develop a web-based tool for electricity producers and consumers to compare renewable and conventional energy technologies to the conventional and advanced nuclear fuel cycle options. The work leveraged data on land use, carbon intensity, unit cost and pricing data for renewables from the open literature. Input data from the Advanced Fuel Cycle Cost Basis report will be used for current and advanced nuclear power systems. We coupled these data to energy flow and lifecycle models for user-selected energy generation and storage types. We will display economic comparisons from the perspectives of generators and consumers, as well as carbon production, land use, and reliability comparisons.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

HydroForecast Long-term: Improving hydropower’s resilience to climate change through accurate climate-scale

With hydrologic patterns and water availability across the globe shifting due to climate change, advancements in hydrologic prediction systems can help significantly reduce the uncertainties that utilities and water supply entities have in their decision making. Understanding and estimating hydrology at the climate scale is critical for managing water resources under changing climate scenarios. This project focuses on integrating state-of-the-art neural network modeling with downscaled climate projections to deliver the reliable water supply projections decades into the future to meet an urgent need from hydropower operators and water utilities. In this Phase 1 DOE SBIR proposal, we developed and validated a theory-guided neural network model, HydroForecast Long-term, for climate-scale hydrology and implemented the model within existing HydroForecast infrastructure. HydroForecast Long-term combines the most accurate streamflow modeling system with a flexible and scalable data architecture to generate water supply projections out to the year 2100. This report illustrates that we have achieved our four objectives: 1) create a prototype of HydroForecast Long-term, building the neural network prediction model, 2) build an automated data input pipeline that processes large amounts of data from the latest global temperature and precipitation climate models; 3) benchmark the accuracy of the hydrologic model over the recent two decades over a large set of diverse basins, and 4) create a set of output visuals and summary metrics informed by customer feedback that connect the data to critical decision points. This work empowers water users to make data-informed decisions supporting a resilient, renewable-powered grid and water system. The results advance the Department of Energy’s mission by addressing critical gaps in water supply planning under climate change.

13 HYDRO ENERGY↗

Data Processing Unit Services Module

The Data Processing Services Module (DPUSM) provides the ability to perform pluggable compression, erasure coding, checksuming and other important file system operations within the Linux kernel. The pluggable provider interface allows for the use of hardware acceleration of those services. In-kernel file systems are then able to use these functions to use these accelerators to perform operations that are normally run on the processor, resulting in improved file system performance. Third parties will register "providers" with the DPUSM to communicate with their respective accelerators. Providers will implement functions with DPUSM API signatures so that the DPUSM can translate the data inputted by users of the DPUSM into data that providers recognize.

Lee, Jason↗

System Modeling of the HTTR and Economic Dispatch Model of the Secondary System

High Temperature Gas-cooled Reactors (HTGRs) can be used for the generation of electricity and their process heat can be used to improve the efficiency of chemical processes such as hydrogen production. The JAEA-operated High-Temperature engineering Test Reactor (HTTR-GT/H2) is exploring using the reactor for electricity and hydrogen production. A RELAP5-3D model of the HTTR-GT/H2 secondary system has been developed using design information. The various components and heat exchangers in the secondary system were modeled and results were compared to the design conditions. The results for the sole-power generation mode were shown to fit the design conditions very well. The largest temperature difference was on the order of 7 K, and the largest pressure difference was on the order of 0.05 MPa. The results for the hydrogen cogeneration mode did not match the design conditions nearly as well. The largest temperature difference was about 39 K and the largest pressure difference was about 0.27 MPa at the compressor outlet. The larger differences for the hydrogen cogeneration mode are attributed to the various complex components and the flow being split in the secondary loop. A transient reduction in heat removal capability of the secondary system was investigated. Reactor temperatures are anticipated to rise as a result. The core reactivity response due to this increase in temperature is investigated and is expected to add negative reactivity to the reactor. An economic dispatch model was developed for a nuclear-driven iodine-sulfur cycle system to determine hydrogen sale prices that would make such a system profitable. The study focuses on the development of the economic model and the role that input data plays on final calculated values. It was found that the input electricity prices, whether using historical data or a host of synthetic time histories, produce significantly different breakeven hydrogen sale prices. As such, great care should be used in these economic dispatch analyses to select reasonable input assumptions.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Geothermal Representation in Power System Models

Power system models generally fail to capture the range of characteristics geothermal resources provide and the value they potentially contribute to decarbonization and reliability of future electricity grids as firm, dispatchable, non-combustion power resources. This study reviews the results of power system modeling efforts to investigate geothermal deployment potential in the United States, including the U.S. DOE GeoVision analysis and ongoing modeling and analysis efforts to support planning and development of future grids with 100% renewable energy in California. Several themes are identified that could be implemented immediately to improve the accuracy of geothermal representation in power system models: consistency of model inputs, modeling of baseload and dispatchable geothermal resources, accurate valuation of grid services, improved representation of capacity factor, use of contemporary LCOE estimates, improved understanding of the evolution of geothermal value, and use of accurate resource potential constraints. Many of the models reviewed produced significantly different amounts of geothermal resource selection - even when modeling the same region and time period. This highlights the variability of inputs and assumptions among models, so creating a consistent set of geothermal inputs is a first step toward more accurate representation of geothermal in models. Research opportunities are identified that could help improve geothermal data inputs in modeling efforts, including analyses of historical data, sensitivity to model inputs, and comparative value of geothermal generators as baseload or dispatchable resources. Outcomes of such research can inform the geothermal community about how best to guide geothermal development toward wider deployment in support of future electricity grids through improved understanding of the evolution of geothermal value over time and the characteristics that contribute to that value.

capacity expansion models↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

High Resolution Siting Suitability of Various Power Plant Technologies

Energy sector planning models determine the aggregate need for new generation, but these models are typically at the state or regional scale and are not equipped to address the wide range of location- and technology-specific issues that are increasingly a factor in power plant siting. These animations demonstrate the aggregate siting suitability of various power plant technology configurations, considering technology-specific factors that can prohibit development. The data presented is from the GRIDCERF (Geospatial Raster Input Data for Capacity Expansion Regional Feasibility) data package. GRIDCERF is a harmonized, open-source geospatial product that can be used to evaluate siting suitability for renewable and non-renewable power plants in the conterminous United States. The animations presented here demonstrate a curated selection of the full suite of technology configurations available. GRIDCERF provides the necessary inputs for models that simulate power plant siting for regional capacity expansion planning such as the Capacity Expansion Regional Feasibility (CERF) model.

Mongird, Kendall [Pacific Northwest National Labor↗

HopPyBar

HopPyBar is a python program to import, analyze, and export split-Hopkinson pressure bar (SHPB, also known as Kolsky bar) data. Traditional analysis offers a black box approach, where input data is converted to analyzed output by performing a series of calculations without user involvement. This program serves as a developmental platform to "white box" the data analysis process. Data streams can be captured (in-situ) to enable advanced or unconventional analyses, statistics, and comparisons. Additionally, the program is geared towards the standardized forms of input and output used at LANL to streamline analysis, but the open nature of the program makes additional input/output schemes straightforward to add. General workflow will import SHPB data in one of a number of formats, identify relevant portions of data signals, and convert to stress-strain-strain rate to show material behavior as a function of dynamic testing.

Morrow, Benjamin↗

A framework for testing soil carbon dynamics post land-use transition in a multisector dynamics model

Soil carbon plays a crucial role in the global carbon cycle. Changes in land use can determine whether carbon is stored or is emitted into the atmosphere as carbon dioxide, which has broad implications for the human and Earth systems. These feedbacks to the carbon cycle and their socio-economic drivers are modelled by many global multisector dynamics models to project future possibilities for the human-Earth system. One notable model of this class is the Global Change Analysis Model (GCAM), which uses a simplified process to model soil organic carbon (SOC) content after land-use transition across 384 land units. While the current GCAM soil carbon framework is based on scientific principles, it has not been tested against experimental data. This work examines rates of SOC change from GCAM input data. Specifically, first order rate constants derived from model inputs were compared to values from two syntheses to assess GCAM’s accuracy. Welch’s t-tests and linear models were used to determine if rate constants were consistent across all tested geographical areas and land-use transition types. While we found that there was general agreement on the direction and magnitude (i.e., rate) of SOC change, the rate constant derived from GCAM and empirical values differed strongly in a subset of specific instances. These results indicate that GCAM’s current SOC dynamics during land use transition successfully capture broad patterns of change in this critical carbon pool, but should be interpreted with caution at finer spatial scales. One potential cause of these discrepancies is our highly aggregated variable, soil timescale, which could be made more granular to improve accuracy. When using economically rooted multisector dynamics models, such as GCAM, it is critical to understand such model limitations for representing specific Earth system processes.

carbon↗

Integrated System Planning: Emerging Software Requirements in the Power Industry

Power system planning software remains fragmented across organizational boundaries, with specialized tools for capacity expansion, production cost modeling, power flow, and dynamic analysis operating on incompatible data models and assumptions. This article argues that the fragmentation is not merely a technical problem but a predictable consequence of Conway's law: software architectures mirror the departmental structures within which they are developed. Regulatory milestones like Federal Energy Regulatory Commission (FERC) Order 888 formalized these divisions, but the roots trace back to the distinct engineering disciplines-mechanical, chemical, and electrical-that staffed generation and transmission planning departments in vertically integrated utilities. As the industry moves toward integrated system planning (ISP) that coordinates generation, transmission, and distribution investment decisions, the software ecosystem must evolve accordingly. We identify five categories of software requirements to enable this transition: coherent data inputs decoupled from individual applications, unified and extensible data schemas, modular component representations that support multiple abstraction levels, lifecycle management of planning datasets, and well-defined application programming interface (API) contracts that separate data exchange from algorithmic control. We examine how these requirements interact with three common workflow patterns-serial gate clearing, sequential multiapplication, and convergence oriented-and discuss the interface design principles each demands. We then outline a vision for platform-based planning architectures where specialized analytical services compose through standardized interfaces and where artificial intelligence (AI)/machine learning (ML) tools augment decision support within a disciplined software infrastructure. The practices proposed here offer a path from today's siloed tool collections toward collaborative planning ecosystems capable of handling the complexity of modern power system transformation.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Probabilistic Reasoner Based on Bayes Risk for Damage Detection in Structural Systems

Structural health monitoring (SHM) systems are used to inform operation of structural systems subject to loads and environments that may affect their integrity. SHM systems rely on continuous monitoring of the structure to determine its health state. These systems are often coupled with a model of the deployed structure to determine the consequences of changes in the system by forecasting the response to future states. These models, which may be thought of as digital twins, need to be updated to reflect the latest state of the structural system. This work makes use of an uncertainty-aware machine learning model that enforces distance preservation of the original input space to determine deviations from the training data input space distributions. This workflow enables domain shift detection to determine whether damage is present in the structure. The uncertainty metrics generated by this network are then used in a Bayes risk framework to design an optimal damage detector given cost and risk considerations. The approach is demonstrated on a computational example with simulated damage.

Najera-Flores, David [ATA Engineering, Inc.]↗

Machine Learning-Based Identification of the Interface Regions for Coupling Local and Nonlocal Models

Local-nonlocal coupling approaches provide a means to combine the computational efficiency of local models and the accuracy of nonlocal models. However, the coupling process can be challenging, requiring expertise to identify the interface between local and nonlocal regions. Here, this study introduces a machine learning-based approach to automatically detect the regions in which the local and nonlocal models should be used. The method uses loading functions evaluated at grid points to decide the model selection at those points. Training of the networks is based on datasets provided by classes of loading functions for which reference coupling configurations are computed using accurate coupled solutions, where accuracy is measured in terms of the relative error between the solution to the coupling approach and the solution to the nonlocal model. We study two approaches that vary in data structure. The first, the full-domain input data approach, uses the entire load vector and outputs a complete label vector, performing a global classification. The second, a window-based approach, processes loads into windows and addresses the problem as a node-wise classification where each window's central point is classified individually. The classification problems are solved via deep learning algorithms based on convolutional neural networks. The performance of these approaches is studied on one-dimensional numerical examples using F1-scores and accuracy metrics. Notably, the windowing approach achieves an accuracy of 0.96 and an F1-score of 0.97, highlighting its potential to automate coupling processes effectively and enhance computational efficiency in material science applications.

97 MATHEMATICS AND COMPUTING↗

ARM Data-Oriented Metrics and Diagnostics Package for Climate Model Evaluation

A Python-based metrics and diagnostics package is currently being developed by the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Infrastructure Team at Lawrence Livermore National Laboratory (LLNL) to facilitate the use of long-term, high-frequency measurements from the ARM Facility in evaluating the regional climate simulation of clouds, radiation, and precipitation. This metrics and diagnostics package computes climatological means of targeted climate model simulation and generates tables and plots for comparing the model simulation with ARM observational data. The Coupled Model Intercomparison Project (CMIP) model data sets are also included in the package to enable model intercomparison as demonstrated in Zhang et al. (2017). The mean of the CMIP model can serve as a reference for individual models. Basic performance metrics are computed to measure the accuracy of mean state and variability of climate models. The evaluated physical quantities include cloud fraction, temperature, relative humidity, cloud liquid water path, total column water vapor, precipitation, sensible and latent heat fluxes, and radiative fluxes, with plan to extend to more fields, such as aerosol and microphysics properties. Process-oriented diagnostics focusing on individual cloud- and precipitation-related phenomena are also being developed for the evaluation and development of specific model physical parameterizations. The version 1.0 package is designed based on data collected at ARM’s Southern Great Plains (SGP) Research Facility, with the plan to extend to other ARM sites. The metrics and diagnostics package is currently built upon standard Python libraries and additional Python packages developed by DOE (such as CDMS and CDAT). The ARM metrics and diagnostic package is available publicly with the hope that it can serve as an easy entry point for climate modelers to compare their models with ARM data. In this report, we first present the input data, which constitutes the core content of the metrics and diagnostics package in section 2, and a user's guide documenting the workflow/structure of the version 1.0 codes, and including step-by-step instruction for running the package in section 3.

54 ENVIRONMENTAL SCIENCES↗

Blockchain based Communication Architectures with Applications to Private Security Networks

Existing communication protocols in high consequence security networks are highly centralized. While this naively makes the controls easier to physically secure, external actors require fewer resources to disrupt the system because there are fewer points in the system can be destroyed or interrupted without the entire system failing. We present a solution to this problem using a proof-of-work-based blockchain implementation built on MultiChain. We construct a test-bed network containing two types of data input: visual imagers and microwave sensor information. These data types are ubiquitous in perimeter intrusion detection security systems and allow a realistic representation of a real-world network architecture. The cameras in this system use an object detection algorithm to nd important targets in the scene. The raw data from the camera and the outputs from the detection algorithm are then placed in a transaction on the distributed ledger. Similarly, microwave data is used to detect relevant events and are placed in a transaction. These transactions are then bundled into blocks and broadcast to the rest of the network using the Bitcoin-based MultiChain protocol. We develop five tests to examine the security metrics of our network. We performed the five security metric test using different sized networks from 7 to 39 nodes to determine how the metrics scale with respect to size. We nd that when compared to a centralized architecture our implementation provides a resiliency increase that is expected from a blockchain-based protocol without slowing the system so much that a human operator would notice. Furthermore, our approach is able to detect tampering in real time. Based on these results, we theorize that security networks in general could use a blockchain-based approach in a meaningful way.

97 MATHEMATICS AND COMPUTING↗

Updates to the ATLAS Data Carousel Project

The High Luminosity upgrade to the LHC (HL-LHC) is expected to deliver scientific data at the multi-exabyte scale. In order to address this unprecedented data storage challenge, the ATLAS experiment launched the Data Carousel project in 2018. Data Carousel is a tape-driven workflow whereby bulk production campaigns with input data resident on tape are executed by staging and promptly processing a sliding window to disk buffer such that only a small fraction of inputs are pinned on disk at any one time. Data Carousel is now in production for ATLAS in Run3. In this paper, we provide updates on recent Data Carousel R&D projects, including data-on-demand and tape smart writing. Data-on-demand removes from disk data that has not been accessed for a predefined period, when users request them, they will be either staged from tape or recreated by following the original production steps. Tape smart writing employs intelligent algorithms for file placement on tape in order to retrieve data back more efficiently, which is our long term strategy to achieve optimal tape usage in Data Carousel.

42 ENGINEERING↗

RODeO (Revenue Operation and Device Optimization Model) [SWR 20-67]

The Revenue, Operation, and Device Optimization (RODeO) model explores optimal system design and operation considering different levels of grid integration, equipment cost, operating limitations, financing, and credits and incentives. RODeO is a price-taker model formulated as a mixed-integer linear programming (MILP) model in the GAMS modeling platform. The objective is to maximizes the net revenue for a collection of equipment at a given site. The equipment includes generators (e.g., gas turbine, steam turbine, solar, wind, hydro, fuel cells, etc.), storage systems (batteries, pumped hydro, gas-fired compressed air energy storage, long-duration systems, hydrogen), and flexible loads (e.g., electric vehicles, electrolyzers, flexible building loads). The input data required by RODeO can be classified into three bins: 1) utility service data, which refers to retail utility rate information (meter cost, energy and demand charges), 2) electricity market data, which include energy and reserve prices, 3) other inputs, which refer to additional electrical demand, product output demand, technological assumptions, financial properties, and operational parameters.

Guerra Fernandez, Omar Jose↗