Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Multi-Parent Clustering Algorithms from Stochastic Grammar Data Models

We introduce a statistical data model and an associated optimization-based clustering algorithm which allows data vectors to belong to zero, one or several "parent" clusters. For each data vector the algorithm makes a discrete decision among these alternatives. Thus, a recursive version of this algorithm would place data clusters in a Directed Acyclic Graph rather than a tree. We test the algorithm with synthetic data generated according to the statistical data model. We also illustrate the algorithm using real data from large-scale gene expression assays.

Mjoisness, Eric↗

Enhanced Modeling of First-Order Plant Equations of Motion for Aeroelastic and Aeroservoelastic Applications

A methodology is described for generating first-order plant equations of motion for aeroelastic and aeroservoelastic applications. The description begins with the process of generating data files representing specialized mode-shapes, such as rigid-body and control surface modes, using both PATRAN and NASTRAN analysis. NASTRAN executes the 146 solution sequence using numerous Direct Matrix Abstraction Program (DMAP) calls to import the mode-shape files and to perform the aeroelastic response analysis. The aeroelastic response analysis calculates and extracts structural frequencies, generalized masses, frequency-dependent generalized aerodynamic force (GAF) coefficients, sensor deflections and load coefficients data as text-formatted data files. The data files are then re-sequenced and re-formatted using a custom written FORTRAN program. The text-formatted data files are stored and coefficients for s-plane equations are fitted to the frequency-dependent GAF coefficients using two Interactions of Structures, Aerodynamics and Controls (ISAC) programs. With tabular files from stored data created by ISAC, MATLAB generates the first-order aeroservoelastic plant equations of motion. These equations include control-surface actuator, turbulence, sensor and load modeling. Altitude varying root-locus plot and PSD plot results for a model of the F-18 aircraft are presented to demonstrate the capability.

Pototzky, Anthony S.↗

Hot-Fire Testing of 5N and 22N HPGP Thrusters

This hot-fire test continues NASA investigation of green propellant technologies for future missions. To show the potential for green propellants to replace some hydrazine systems in future spacecraft, NASA Marshall Space Flight Center (MSFC) is continuing to embark on hot-fire test campaigns with various green propellant blends.NASA completed hot-fire testing of 5N and 22N HPGP thrusters at the Marshall Space Flight Center’s Component Development Area altitude test stand in April 2015. Both thrusters are ground test articles and not flight ready units, but are representative of potential flight hardware with a known path towards flight application. The purpose of the 5N testing was to perform facility check-outs and generate a small set of data for comparison to ECAPS and Orbital ATK data sets. The 5N thruster performed as expected with thrust and propellant flow-rate data generated that are similar to previous testing at Orbital ATK. Immediately following the 5N testing, and using the same facility, the 22N testing was conducted on the same test stand with the purpose of demonstrating the 22N performance. The results of 22N testing indicate it performed as expected.The results of the hot-fire testing are presented in this paper and presentation.

Burnside, Christopher G.↗

Transitioning the NASA SLR Network to Event Timing Mode for Reduced Systematics, Improved Stability and Data Precision

NASA's legacy Satellite Laser Ranging (SLR) network produces about one-third of the global SLR data to support spacegeodesy. This network of globally distributed stations has been using Time Interval Units (TIU) for range measurements for thelast 25 + years. To improve the reliability of the SLR network and satisfy the need for stable millimeter precision data, a phasedreplacement of the TIUs in the network with picosecond-precise Event Timer Modules was initiated in 2015. This schemeallowed the time of flight and laser transmit epoch measurement to one picosecond resolution. For a network with globalscientific impact, transitioning to a new data generation metrological scheme requires significant data scrutiny and long-termscience data validation. Any long-term testing/measurement has the potential to interrupt the station's daily operational dataflow to the International Laser Ranging Service (ILRS) as the station under test will have to put its test data into quarantine.We have demonstrated a very effective way to test and implement the new device without removing the old hardware andwithout the need for the orbit analysis. This operationally noninvasive scheme performed concurrent test measurements enablinguninterrupted operational data flow to the users, while allowing simultaneous test data capture for short- and long-termsystematics and stability analysis. Extensive analysis of the test data was performed by the NASA SLR engineering team andthe ILRS Analysis Standing Committee, to uncover biases and any dependencies on the satellite ranges (for nonlinear scaleissues). Multi-ETM comparison was also performed at two of the SLR stations through the interchange of hardware to establishthe inter-device range biases and stability. Such benchmarked hardware was subsequently sent to the remaining stationsto allow traceability and normalize the network performance. The range bias intercomparison performed using the multiyearSLR data analysis agreed well with the engineering changes, thus validating the approach to flush out station-specific rangingsystematics affecting precise orbit determination. Such an improvement and rebalancing of the current network will allowan orderly transition of the current NASA SLR network operating at a maximum rate of 10 Hz to the NASA next generationSpace Geodesy Satellite Laser Ranging (SGSLR) network operating at 2 kHz (McGarry et al. in J Geod, 2018. https ://doi.org/10.1007/s0019 0-018-1191-6; Merkowitz et al. in J Geod, 2018. https ://doi.org/10.1007/s0019 0-018-1204-5).

Varghese, Thomas↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Temporal Study 2022-2024: Sample-Based Surface Water Dissolved Inorganic Carbon, Dissolved Organic Carbon, Total Nitrogen, Stable Isotopes, and Total Suspended Solids from across Multiple Watersheds in the Yakima River Basin, Washington, USA

This dataset supports a broader study examining the drivers of temporal variability in sediment respiration rates in the Yakima River Basin. The dataset provides geochemistry data generated from samples collected at bi-weekly or monthly intervals at six sites across the Yakima River Basin in Washington, USA. Sample and sensor data from previous years (2021-2022) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1898912 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1892054, respectively. Related sensor data from 2022-2024 will be published separately. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) dissolved inorganic carbon (DIC) and averages; (6) dissolved organic carbon (DOC; reported as non-purgeable organic carbon; NPOC) and averages; (7) total dissolved nitrogen (TN) and averages; (8) total suspended solids (TSS); (9) stable isotopes; (10) surface water sampling protocol; (11) sensor protocol; (12) methods codes; and (13) international generic sample number (IGSN) mapping file. All files are .csv or .pdf. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. For data and scripts associated with "Shifts in rain-snow partitioning drive faster water transit times in the US Pacific Northwest" (Butler et al., 2026), go to https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3025481

18-O↗

Laboratory time series moisture manipulative experiment from sediment across San Antonio, Texas: time series aerobic respiration and geochemistry

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration. The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS Allison Veach collaboration (AV1). The data package associated with the AV1 study is available at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2529428. AV1 sampling occurred across 7 perennial and 7 intermittent streams in San Antonio, Texas. Each stream/site was visited both in summer during base flow (July-September 2023) and winter during peak flow (January-February 2024). This study uses subsamples from a subset of AV1 samples. The original field samples were labeled as AV1_###. Subsequent subsamples for this study were labeled as EV_###. The labels from the field samples and the EV subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EV_001 is a subsample from AV1_001). See the critical details section below for more details on sample naming. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) effect size; (2) iron (II); (3) gravimetric moisture; (4) respiration rates; (5) raw dissolved oxygen values and plots; (6) specific conductance; (7) pH; (8) temperature; (9) a summary containing mean, median, and standard deviation values of each data type for each treatment (wet and dry); and (10) methods codes. All files are .csv or.pdf.

54 ENVIRONMENTAL SCIENCES↗

Data Analysis with Graphical Models: Software Tools

Probabilistic graphical models (directed and undirected Markov fields, and combined in chain graphs) are used widely in expert systems, image processing and other areas as a framework for representing and reasoning with probabilities. They come with corresponding algorithms for performing probabilistic inference. This paper discusses an extension to these models by Spiegelhalter and Gilks, plates, used to graphically model the notion of a sample. This offers a graphical specification language for representing data analysis problems. When combined with general methods for statistical inference, this also offers a unifying framework for prototyping and/or generating data analysis algorithms from graphical specifications. This paper outlines the framework and then presents some basic tools for the task: a graphical version of the Pitman-Koopman Theorem for the exponential family, problem decomposition, and the calculation of exact Bayes factors. Other tools already developed, such as automatic differentiation, Gibbs sampling, and use of the EM algorithm, make this a broad basis for the generation of data analysis software.

Buntine, Wray L.↗

Multi-agent voltage control in distribution systems using GAN-DRL-based approach

Active distribution grids can experience voltage fluctuations and violations due to the high penetration of variable distributed energy resources (DERs). These problems might occur because of the uncertain and variable generation natures of these resources, especially solar photovoltaic resources, during panel shadowing scenarios. Volt-VAR control (VVC) is an efficient method that controls the reactive power set-points of the inverters to regulate the voltage of distribution grids. Although several VVC approaches have been proposed recently, the performance of these approaches degrades significantly if behind-the-meter solar generation data are unobservable/missing. Therefore, it is necessary to impute missing/unobservable PV data accurately to be utilized in VVC approaches. Further, this paper proposes a model-free, data-driven, centrally trained, and decentrally executed multi-agent deep reinforcement learning-based VVC architecture to regulate the voltage of distribution networks. A generative adversarial network (GAN) is incorporated to impute the unobservable PV data accurately, which improves the performance of the proposed control architecture. The proposed multi-agent-soft-actor–critic algorithm (MASAC)-based VVC technique utilizes the actual PV dataset as well as the imputed dataset from the GAN framework to learn the optimal coordinated control policy for controlling the optimal reactive power set-points of PV inverters. The effectiveness of the proposed approach is analyzed on a modified IEEE 34-bus test case with added PV inverters. The results are compared and analyzed with a base case model with no VVC and VVC with a local droop control approach, genetic algorithm optimization, and a centralized soft actor–critic-based approach. Moreover, the performance of the proposed approach is compared with that of a multi-agent VVC framework without using the PV generation data and load information as the system state. The results illustrate that the proposed method with more state input improves the voltage profile and reduces the power loss of the network across various loading and PV generation scenarios.

14 SOLAR ENERGY↗

Definition of common support equipment and space station interface requirements for IOC model technology experiments

A study was conducted to identify the common support equipment and Space Station interface requirements for the IOC (initial operating capabilities) model technology experiments. In particular, each principal investigator for the proposed model technology experiment was contacted and visited for technical understanding and support for the generation of the detailed technical backup data required for completion of this study. Based on the data generated, a strong case can be made for a dedicated technology experiment command and control work station consisting of a command keyboard, cathode ray tube, data processing and storage, and an alert/annunciator panel located in the pressurized laboratory.

Russell, Richard A.↗

Between a Map and a Data Rod

A Digital Divide has long stood between how NASA and other satellite-derived data are typically archived (time-step arrays or maps) and how hydrology and other point-time series oriented communities prefer to access those data. In essence, the desired method of data access is orthogonal to the way the data are archived. Our approach to bridging the Divide is part of a larger NASA-supported data rods project to enhance access to and use of NASA and other data by the Consortium of Universities for the Advancement of Hydrologic Science, Inc. (CUAHSI) Hydrologic Information System (HIS) and the larger hydrology community. Our main objective was to determine a way to reorganize data that is optimal for these communities. Two related objectives were to optimally reorganize data in a way that (1) is operational and fits in and leverages the existing Goddard Earth Sciences Data and Information Services Center (GES DISC) operational environment and (2) addresses the scaling up of data sets available as time series from those archived at the GES DISC to potentially include those from other Earth Observing System Data and Information System (EOSDIS) data archives. Through several prototype efforts and lessons learned, we arrived at a non-database solution that satisfied our objectivesconstraints. We describe, in this presentation, how we implemented the operational production of pre-generated data rods and, considering the tradeoffs between length of time series (or number of time steps), resources needed, and performance, how we implemented the operational production of on-the-fly (virtual) data rods. For the virtual data rods, we leveraged a number of existing resources, including the NASA Giovanni Cache and NetCDF Operators (NCO) and used data cubes processed in parallel. Our current benchmark performance for virtual generation of data rods is about a years worth of time series for hourly data (9,000 time steps) in 90 seconds. Our approach is a specific implementation of the general optimal strategy of reorganizing data to match the desired means of access. Results from our project have already significantly extended NASA data to the large and important hydrology user community that has been, heretofore, mostly unable to easily access and use NASA data.

machine learning↗

Machine Learning-Enhanced Multiphase CFD for Carbon Capture Modeling Run Data

Repository for the data generated as part of the 2023-2024 ALCC project "Machine Learning-Enhanced Multiphase CFD for Carbon Capture Modeling." The data was generated with MFIX-Exa's CFD-DEM model. The problem of interest is gravity driven, particle-laden, gas-solid flow in a triply-periodic domain of length 2048 particle diameters with an aspect ratio of 4. The mean particle concentration ranges from 1% to 40% and the Archimedes number ranges from 18 to 90. The particle-to-fluid density ratio, particle-particle restitution and friction coefficients and domain aspect ratio are held constant at values of 1000, 0.9, 0.25 and 4, respectively. This research used resources of the National Energy Research Scientific Computing Center, a DOE Office of Science User Facility supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231 using NERSC award ALCC-ERCAP0025948.

AMReX↗

Pyrogenic Organic Matter Laboratory Experiment: Aerobic Respiration and Geochemistry from Variably Inundated Stream Sediments (v3)

This dataset supports a broader study examining the effects of variable inundation and pyrogenic organic matter on ecosystem respiration. The dataset provides data generated from a laboratory batch experiment investigating the interaction between variable inundation conditions (wet and dry sediment) and pyrogenic organic matter (burned and unburned treatments). The contents include time series dissolved oxygen, sediment geochemistry data, and field metadata (including qualitative information on instream and river corridor characteristics). This data package was originally published in November 2025. It was updated in April 2026 (v2; new and modified files) and May 2026 (v3; modified files). See the change history section in the readme for more details For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) international generic sample number (IGSN) mapping file; (5) readme; (6) field protocol; (7) sample name metadata; (8) an environmental context picture for the dry and inundated sampling locations; and (9) a subfolder with sample data from the sediment incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) gravimetric moisture; (4) partial pressure and production rates of carbon dioxide, methane, and nitrous oxide; (5) field wet sediment mass, dry sediment mass, water mass, and field wet sediment volume in incubation and sediment NPOC/TN vials; (6) methods codes; (7) respiration rates, pH, and temperature from after the incubation, raw time series dissolved oxygen and temperature, and a subfolder containing associated plots and scripts; (8) ions; (9) FTICR-MS methods; and (10) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains the CoreMS processed data and three subfolders, one containing the .xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

Aggregation Tool to Create Curated Data albums to Support Disaster Recovery and Response

Despite advances in science and technology of prediction and simulation of natural hazards, losses incurred due to natural disasters keep growing every year. Natural disasters cause more economic losses as compared to anthropogenic disasters. Economic losses due to natural hazards are estimated to be around $6-$10 billion dollars annually for the U.S. and this number keeps increasing every year. This increase has been attributed to population growth and migration to more hazard prone locations such as coasts. As this trend continues, in concert with shifts in weather patterns caused by climate change, it is anticipated that losses associated with natural disasters will keep growing substantially. One of challenges disaster response and recovery analysts face is to quickly find, access and utilize a vast variety of relevant geospatial data collected by different federal agencies such as DoD, NASA, NOAA, EPA, USGS etc. Some examples of these data sets include high spatio-temporal resolution multi/hyperspectral satellite imagery, model prediction outputs from weather models, latest radar scans, measurements from an array of sensor networks such as Integrated Ocean Observing System etc. More often analysts may be familiar with limited, but specific datasets and are often unaware of or unfamiliar with a large quantity of other useful resources. Finding airborne or satellite data useful to a natural disaster event often requires a time consuming search through web pages and data archives. Additional information related to damages, deaths, and injuries requires extensive online searches for news reports and official report summaries. An analyst must also sift through vast amounts of potentially useful digital information captured by the general public such as geo-tagged photos, videos and real time damage updates within twitter feeds. Collecting and aggregating these information fragments can provide useful information in assessing damage in real time and help direct recovery efforts. The search process for the analyst could be made much more efficient and productive if a tool could go beyond a typical search engine and provide not just links to web sites but actual links to specific data relevant to the natural disaster, parse unstructured reports for useful information nuggets, as well as gather other related reports, summaries, news stories, and images. This presentation will describe a semantic aggregation tool developed to address similar problem for Earth Science researchers. This tool provides automated curation, and creates "Data Albums" to support case studies. The generated "Data Albums" are compiled collections of information related to a specific science topic or event, containing links to relevant data files (granules) from different instruments; tools and services for visualization and analysis; information about the event contained in news reports, and images or videos to supplement research analysis. An ontology-based relevancy-ranking algorithm drives the curation of relevant data sets for a given event. This tool is now being used to generate a catalog of Hurricane Case Studies at Global Hydrology Resource Center (GHRC), one of NASA's Distribute Active Archive Centers. Another instance of the Data Albums tool is currently being created in collaboration with NASA/MSFC's SPoRT Center, which conducts research on unique NASA products and capabilities that can be transitioned to the operational community to solve forecast problems. This new instance focuses on severe weather to support SPoRT researchers in their model evaluation studies

Ramachandran, Rahul↗

Convolutional Variational Autoencoder-based Unsupervised Learning for Power Systems Faults

Classification of power system event data is a growing need, particularly where non-protective relaying-based sensors are used to monitor grid performance. Given the high burden of obtaining event data with appropriate labeling, an unsupervised approach is highly valuable. This approach enables using event data without labeling, which is far easier to obtain. This paper presents an unsupervised learning method to classify and label transients observed in the distribution grid. A Convolutional Variational Autoencoder (CVAE) was developed for this purpose. We demonstrate the efficacy of our approach using the transient data generated from the simulations. The simulation data is used to train the CVAE that identifies different faults as different clusters in the latent space. The clusters are then used as the foundation model to categorize the real-world data.

Alam, Maksudul↗

Computer/computer interface

System synchronizes data transfer between two computers by generating data strobe pulses when computers are ready for data transfer. In addition, interface filters noise by sampling.

Anderson, T. O.↗

Low strain creep and aging of aluminum alloy 2219-T87 sheet

The constant load creep and isothermal aging characteristics of aluminum alloy 2219-T87 sheet have been studied experimentally and analytically in the temperature range 250 to 650 F at stress levels between 2.9 and 4.0 ksi (20 to 283 MPa). Testing variables were closely and automatically monitored. The data generated agree somewhat with the literature data base at lower temperatures, but above 500 F, discrepancies of greater than an order of magnitude in the time to 1% creep strain occur. Good correlation was found with the Larson-Miller parameter as modeled by a second-order polynomial in stress. Constitutive equations for time to 0.1%, 0.2%, 0.5%, and 1.0% creep are given. Information on residual mechanical properties and electrical conductivity is also provided.

Navrotski, G.↗

An Interpolation Method for Obtaining Thermodynamic Properties Near Saturated Liquid and Saturated Vapor Lines

The availability and proper utilization of fluid properties is of fundamental importance in the process of mathematical modeling of propulsion systems. Real fluid properties provide the bridge between the realm of pure analytiis and empirical reality. The two most common approaches used to formulate thermodynamic properties of pure substances are fundamental (or characteristic) equations of state (Helmholtz and Gibbs functions) and a piecemeal approach that is described, for example, in Adebiyi and Russell (1992). This paper neither presents a different method to formulate thermodynamic properties of pure substances nor validates the aforementioned approaches. Rather its purpose is to present a method to be used to facilitate the accurate interpretation of fluid thermodynamic property data generated by existing property packages. There are two parts to this paper. The first part of the paper shows how efficient and usable property tables were generated, with the minimum number of data points, using an aerospace industry standard property package (based on fundamental equations of state approach). The second part describes an innovative interpolation technique that has been developed to properly obtain thermodynamic properties near the saturated liquid and saturated vapor lines.

Nguyen, Huy H.↗