Engineering PapersSearch

SEARCH · Engineering Papers

Results for “DATA BASES”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network

Data-based filtered dissipation rate modelling for multi-modal turbulent combustion: evaluating a priori model generalizability

Manifold-based models offer a computationally efficient alternative to directly transporting the thermochemical state in computational simulations of turbulent reacting flows, projecting the high-dimensional thermochemical state-space onto a low-dimensional manifold. Recent efforts have yielded a manifold-based model applicable to multi-modal combustion, enabling reconstruction of the thermochemical state from solutions to two-dimensional manifold equations in mixture fraction and generalized progress variable that are parameterised by three scalar dissipation rates. In coarse-grained simulations such as Large Eddy Simulation (LES), closure of the multi-modal manifold equations and subfilter variances/covariance requires closure of three filtered scalar dissipation rates. Here, the present work adopts a data-based approach, providing closure for the three filtered scalar dissipation rates via deep neural networks (DNNs). High-fidelity datasets corresponding to an autoigniting n-dodecane jet flame and a bluff body swirl-stabilized confined lifted spray flame of two aviation fuels (Jet-A and C1) with different ignition propensities are leveraged to generate training data that spans a diverse range of thermodynamic conditions and combustion modes, including low- and high-temperature ignition regimes in addition to premixed and nonpremixed behaviour. A final DNN model is trained to enforce inherent physical constraints by learning nonlinear functional transformations of the three filtered scalar dissipation rates. The generalizability of this constrained DNN model is demonstrated a priori via conditional statistics evaluated on the lifted spray flame with C1–a configuration that had not been included in the training data. Excellent DNN agreement with conditional DNS statistics is observed, and integrated gradients are computed to identify the most sensitive input variables. The similarity of the marginal PDFs of the most informative input variables and outputs across configurations are quantified via the Wasserstein metric, demonstrating that data-based models may successfully generalize to unseen parametric conditions so long as the most informative input variables share similar distributions across training and testing datasets.

Data-based modelling

In-pixel integration of signal processing and AI/ML based data filtering for particle tracking detectors

We present the first physical realization of in-pixel signal processing with integrated AI-based data filtering for particle tracking detectors. Building on prior work that demonstrated a physics-motivated edge-AI algorithm suitable for ASIC implementation, this work marks a significant milestone toward intelligent silicon trackers. Our prototype readout chip performs real-time data reduction at the sensor level while meeting stringent requirements on power, area, and latency. The chip is taped-out in 28nm TSMC CMOS bulk process, which has been shown to have sufficient radiation hardness for particle experiments. This development represents a key step toward enabling fully on-detector edge AI, with broad implications for data throughput and discovery potential in high-rate, high-radiation environments such as the High-Luminosity LHC.

Parpillon, Benjamin [Fermilab; Illinois U., Chicag

Radiative impact of record-breaking wildfires from integrated ground-based data

The radiative effects of wildfires have been traditionally estimated by models using radiative transfer calculations. Assessment of model-predicted radiative effects commonly involves information on observation-based aerosol optical properties. However, lack or incompleteness of this information for dense plumes generated by intense wildfires reduces substantially the applicability of this assessment. Here we introduce a novel method that provides additional observational constraints for such assessments using widely available ground-based measurements of shortwave and spectrally resolved irradiances and aerosol optical depth (AOD) in the visible and near-infrared spectral ranges. We apply our method to quantify the radiative impact of the record-breaking wildfires that occurred in the Western US in September 2020. For our quantification we use integrated ground-based data collected at the Atmospheric Measurements Laboratory in Richland, Washington, USA with a location frequently downwind of wildfires in the Western US. We demonstrate that remarkably dense plumes generated by these wildfires strongly reduced the solar surface irradiance (up to 70% or 450 Wm -2 for total shortwave flux) and almost completely masked the sun from view due to extremely large AOD (above 10 at 500 nm wavelength). We also demonstrate that the plume-induced radiative impact is comparable in magnitude with those produced by a violent volcano eruption occurred in the Western US in 1980 and continental cumuli.

54 ENVIRONMENTAL SCIENCES

Data Management in the Continuum: Cross-facility Object-based Data Transfers

Scientific workflows are evolving from relying on a monolithic storage subsystem at a single High-Performance Computing (HPC) facility to using geographically distributed file systems, repositories, and cloud storage. As a result, storing, accessing, transferring, and managing scientific data have become highly complex and prone to performance inefficiencies. This paper delves into these challenges by exploring an optimized end-to-end interface designed to seamlessly connect various local and remote storage systems, enabling efficient data movement of objects across HPC–Cloud and HPC–HPC environments. We showcase this capability through an object-focused data management runtime system, discuss the effects of relaxed consistency semantics in distributed object scenarios, and illustrate its application in an earthquake simulation workflow. Besides reducing the amount of data by selectively transferring regions of interest, our facility-local results achieved a speedup of 45 × over an optimized HDF5 usage and 15 × over the HDF5 with caching by using the new interface in PDC-XF.

Bez, Jean Luca

CO2 Plume Imaging with Accelerated Deep Learning-based Data Assimilation Considering Multiple Realizations: Application to the Illinois Basin-Decatur Carbon Sequestration Project

We propose a fast and efficient deep learning workflow for near real-time data assimilation, forecasting and visualization of CO2 plume evolution in saline aquifer and demonstrate its application at a field site. Unlike the previous work, this study incorporates the impact of spatial heterogeneity using multiple realizations. In the proposed workflow, a neural network model utilizes available monitoring data such as downhole pressure measurements as input and predicts the propagating pressure ‘front’ using the diffusive time of flight (DTOF) map which is considered as representative reservoir image of the flow field. The DTOF is the arrival time of pressure front propagation, which can be computed by the Fast Marching Method rapidly without flow simulations. Reservoir model calibration can be implemented by selecting the training data samples that describe the predicted DTOF map based on observed data. The power and efficacy of our workflow is demonstrated by application to the Illinois Basin-Decatur Project.

CO2 plume imaging

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY

ML-based Data Assimilation and History Matching: Application to the IBDP CCS Project

It is crucial to monitor the CO2 plume effectively throughout the life cycle of a geologic CO2 sequestration project to ensure safety and storage efficiency. However, the computational cost of existing data assimilation methods can be prohibitively expensive due to the complex physics with multi-component non-isothermal simulation and high dimensionality of large-scale reservoir models. We address this challenge by proposing an accelerated deep learning-based workflow for model calibration and prediction of CO2 plume evolution in the reservoir.The power and efficacy of our workflow is demonstrated by application to the Illinois Basin-Decatur Project (IBDP), a large-scale CO2 storage test in saline aquifer. The data assimilation process is implemented rapidly by the proposed workflow with given field measurements including distributed pressure and temperature sensing (DTS) data at an injection and a monitoring well. CO2 plume evolution is predicted by running the simulations of the calibrated reservoir models.

Nagao, Masahiro

A Framework to Demonstrate a DNP3 Interface With a CIM-Based Data Integration Platform: Preprint

The contemporary electrical grid is characterized by its complexity and abundance of data. A control-rich environment supported by information and communication technologies within an Advanced Distribution Management System (ADMS) presents a viable and cost-effective option for utility companies aiming to implement advanced real-time analytical schemes for monitoring and remotely controlling distribution feeders. Modular platform-based approaches to distribution operations require a structured framework for acquiring field device measurements, performing analytics, converting the setpoint to the correct protocol, and sending it on the appropriate communications network to the field devices. We present the development and deployment of an application service to integrate an open-source standardsbased platform with an ADMS test bed with field devices using the Distributed Network Protocol (DNP3) for data exchange. The step-by-step procedure for establishing the DNP3-Master service on an open-source distribution platform is outlined, comprehensively explaining the Master setup process. Moreover, sample use case results highlight the capabilities of the DNP3- Master service setup. Results demonstrate the scalability and configurability of the DNP3-Master service, making it adaptable for integration with other relevant applications, thus providing potential opportunities for real-world field trials and real-time assessments.

ADMS

5-minute Wind Power Data based on WFIP2 WRF Simulation

The second Wind Forecast Improvement Project (WFIP2) was a public-private partnership funded by the U.S. Department of Energy and NOAA, aimed at enhancing the forecast skill of numerical weather prediction models for turbine-height winds in regions with complex terrain. An 18-month Weather Research and Forecasting (WRF) model simulation was conducted over the Pacific Northwest, with model outputs validated against observational data collected during WFIP2. Simulated wind speeds were used to estimate wind power generation using reV (the Renewable Energy Potential Model developed by NREL) at ten wind project sites. Two sets of results were produced: one using wind speeds extracted from the model grid cell at the project centroid, and another using wind speeds from the actual turbine locations. For each dataset, power output was calculated using both actual turbine-specific power curves and nine generic power curves to convert wind speed into power.

17 WIND ENERGY

LogPath: Log data based energy consumption analysis enabling electric vehicle path optimization

Vehicle navigation and path optimization require a more meticulous approach when it deals with EVs (electric vehicles) and SDVs (software-defined vehicles), due to lengthy charging times and the lack of charging infrastructure. Long-distance freight EV trucking needs path guidance with accurate energy consumption estimates to prevent charging-related failures. We developed a novel energy consumption estimation approach that only uses battery log data to extract major vehicle parameters to increase EV navigation accuracy without additional sensors. This is enabled by extracting multiple drive modes from the log data for analysis. The system provides 1) routes, 2) charge locations, 3) charging times, and 4) optimal vehicle speeds that guarantee the shortest travel time. Here we successfully validated the system using log data collected from an EV and Tesla's Supercharging map in the US and compared it with the commercially available navigation system, Tesla's trip planner, whose capabilities solely include charging time and routing.

EV (Electric vehicles) navigation

The National Climate Data Base (NCDB): A Bias-Corrected High-Resolution Climate Dataset

Assessing renewable energy resources under future climate scenarios has been highlighted in recent years to analyze and understand potential impacts of future change in renewable generation on the power sector. Solar energy is well-known as the most plentiful among various renewable resources and usually converted to electricity using photovoltaics (PV) technologies, and the global deployment of PV technology has increased rapidly in recent decades. In this study, we develop a statistical technique to downscale the future projection of solar irradiance for PV energy-related applications. A set of Regional Climate Model (RCM)-based projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) are used as inputs to statistical methods to generate high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). The main steps of the statistical downscaling method include (1) regridding RCM output (0.22 degree and daily resolutions) to handle the modeled-observed data sets on a common grid, (2) correcting bias of RCM GHI using satellite-derived observation, and (3) implementing temporal and spatial downscaling to generate GHI at 8-km and hourly resolution. Basically, complex physical processes and interactions between solar radiation and various atmospheric constituents lead solar irradiance to be highly variable and uncertain. Underrepresentation of clouds from the RCM parameterizations is the main source of error and uncertainty in modeling solar irradiance. Thus, we adapt and use the high-quality satellite-derived data from the National Solar Radiation Database (NSRDB) to analyze the bias and error of RCM GHI as well as estimate the statistical parameters for spatial and temporal downscaling. This presentation will summarize the comprehensive analysis conducted to produce and assess the results under two climate scenarios (RCP4.5 and RCP8.5). We will also present a detailed validation demonstrating the strengths of the proposed downscaling method and future extension of this research.

climate data

Towards AI Based Data Classification for Decision Making During Testing

During the development of high-consequence items, test systems should be capable of differentiating between test failures resulting from narrowly missing requirements versus those indicating potentially catastrophic faults. In many instances, classifying the data corresponds to simply identifying whether measured waveforms have approximately the anticipated shape. Cast in this light, the problem reduces to converting raw data into a form optimal for use with neural network classifiers. This manuscript investigates different means of representing raw data for image classification. Raw data plots and Short Time Fourier Transform (STFT) spectrograms are classified by both custom built, small-scale, Convolution Neural Networks (CNN) and open-source, multi-million parameter, pre-trained deep CNNs. In the case of time varying frequency content, the STFTs provide images with greater detail and can be accurately classified with simpler networks. This requires less memory and runs faster than classifying the raw data using the more sophisticated options—making STFTs optimal for applications with memory constraints. STFTs are not a panacea. In some cases the time-domain signal contains useful information that should not be discarded. Rather than using raw data or STFTs, the images can be constructed from both by using red and green channels of an RGB image to visualize the real and imaginary components of the transform, with the raw data occupying the blue channel.

97 MATHEMATICS AND COMPUTING