Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Large Dataset Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Unsteady PSP in the NASA Transonic Dynamics Tunnel

For the first time, unsteady pressure sensitive paint (uPSP) has been applied in the NASA Langley Transonic Dynamics Tunnel. Obtaining global surface pressure measurements using the uPSP technique required the development of a new paint formulation for use in the low oxygen heavy gas atmosphere, as well as environmental enclosures to protect sensitive electro-optical components from the high temperature, low pressure environment present during tunnel operation. A high-speed datalink connecting the wind tunnel to Langley’s local high performance compute resource was also established for this test to enable near real time processing of the large datasets that were obtained throughout the campaign. A high-speed lifetime measurement technique was also utilized to yield steady state surface pressures at each condition using the same equipment that was used to provide unsteady measurements. Important metrics such as pressure time histories and power spectral density are compared against traditional unsteady pressure point measurements, and more advanced data products such as dynamic mode decomposition are also explored to provide insight into the underlying flow phenomena.

Daniel T. Reese

EXCLUSIVE NEUTRAL PION ELECTROPRODUCTION CROSS SECTION MEASUREMENTSWITHANEUTRALPARTICLE SPECTROMETER

Deep Virtual Compton Scattering (DVCS), the exclusive electron-proton scattering process ep ¿e'p'¿, provides access to generalized parton distributions (GPDs), which correlate information about the longitudinal momentum and transverse spatial structure of quarks inside the nucleon. Experiment E12-13-010 in Hall C at Jefferson Lab was designed to take high-precision measurements of the DVCS cross section over an extended kinematic range using the newly commissioned Neutral Particle Spectrometer (NPS). The NPS features a high-resolution electromagnetic calorimeter and a streaming data acquisition system optimized for operation at high luminosities. This thesis presents the detector and analysis work carried out to support the NPS DVCS program. In particular, it focuses on the hardware design, calibration, and performance of the calorimeter. A development of a waveform reconstruction analysis of the calorimeter signals enabled improved extraction of pulse amplitudes and times. The waveform analysis was also extended to operate in a multithreaded environment, substantially reducing processing time for large datasets. Analysis of exclusive neutral pion electroproduction events in the calorimeter gives a strong validation of the calorimeter’s performance and resolution. Together these developments establish a foundation for future analyses and extraction of the DVCS cross section and its use in constraining the GPDs.

Kerver, Mitchell [Old Dominion Univ., Norfolk, VA

BAMCensus (The Behavior and Advanced Mobility Census Dataset Aggregator) [SWR-25-120]

This software is a high-performance tool developed in Rust for downloading and processing large-scale geospatial datasets, specifically focusing on US Census data. It is designed to address scaling limitations found in existing tools, such as R's [tidycensus](https://walker-data.com/tidycensus/), by providing performant streaming dataset JOIN operations between various US Census datasets (like ACS and LEHD) and their corresponding geometries stored on the TIGER/Lines web server. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. Its primary motivation stems from the need for a high-performance solution to combine spatial datasets with graph traversals within the context of mobility analysis tooling being developed at NREL's Behavior and Advanced Mobility (BAM) group.

Fitzgerald, Robert [National Renewable Energy Labo

Machine Learning Atom Probe Tomography Tool For Automatic And Fast Clustering

The software uses a YOLO11 segmentation model trained on synthetic data to analyze APT datasets. The workflow operates as follows: 1. Data Slicing: The APT dataset is divided into multiple 2D cross-sections of a specified thickness. 2. Segmentation: The model identifies point-dense regions within each 2D slice. 3. 3D Reconstruction: Detected regions (masks) from all slices are combined and reconstructed back into the original 3D space, forming clusters. The integration with HPC resources enables the software to process large-scale APT datasets efficiently. This combination of automation and scalability reduces manual intervention, improves reproducibility, and accelerates the clustering workflow.

Tang, Yalei [Idaho National Laboratory (INL), Idah

GeoNEX: A Cloud Gateway for Near Real-time Processing of Geostationary Satellite Products

The emergence of a new generation of geostationary satellite sensors provides land andatmosphere monitoring capabilities similar to MODIS and VIIRS with far greater temporal resolution (5-15 minutes). However, processing such large volume, highly dynamic datasets requires computing capabilities that (1) better support data access and knowledge discovery for scientists; (2) provide resources to enable real-time processing for emergency response (wildfire, smoke, dust, etc.); and (3) provide reliable and scalable services for the broader user community. This paper presents an implementation of GeoNEX (Geostationary NASA-NOAA Earth Exchange) services that integrate scientific algorithms with Amazon Web Services (AWS) to provide near realtime monitoring (~5 minute latency) capability in a hybrid cloud-computing environment. It offers a user-friendly, manageable and extendable interface and benefits from the scalability provided by Amazon Web Services. Four use cases are presented to illustrate how to (1) search and access geostationary data; (2) configure computing infrastructure to enable near real-time processing; (3) disseminate and utilize research results, visualizations, and animations to concurrent users; and (4) use a Jupyter Notebook-like interface for data exploration and rapid prototyping. As an example of (3), the Wildfire Automated Biomass Burning Algorithm (WF_ABBA) was implemented on GOES-16 and -17 data to produce an active fire map every 5 minutes over the conterminous US. Details of the implementation strategies, architectures, and challenges of the use cases are discussed.

GeoNEX

Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud

Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.

Lynnes, Christopher

Tracking Provenance of Earth Science Data

Tremendous volumes of data have been captured, archived and analyzed. Sensors, algorithms and processing systems for transforming and analyzing the data are evolving over time. Web Portals and Services can create transient data sets on-demand. Data are transferred from organization to organization with additional transformations at every stage. Provenance in this context refers to the source of data and a record of the process that led to its current state. It encompasses the documentation of a variety of artifacts related to particular data. Provenance is important for understanding and using scientific datasets, and critical for independent confirmation of scientific results. Managing provenance throughout scientific data processing has gained interest lately and there are a variety of approaches. Large scale scientific datasets consisting of thousands to millions of individual data files and processes offer particular challenges. This paper uses the analogy of art history provenance to explore some of the concerns of applying provenance tracking to earth science data. It also illustrates some of the provenance issues with examples drawn from the Ozone Monitoring Instrument (OMI) Data Processing System (OMIDAPS) run at NASA's Goddard Space Flight Center by the first author.

Tilmes, Curt

Automatic detection of microlensing events in the galactic bulge using machine learning techniques

The Wide Field Infrared Survey Telescope (WFIRST) is a NASA flagship mission scheduled to launch in mid-2020, with more than one year of its lifetime dedicated to microlensing survey. The survey is to discover thousands of exoplanets near or beyond the snowline via their microlensing light curve signatures, enabling a Kepler-like statistical analysis of planets at ~1-10 AU from their host stars and potentially revolutionizing our understanding of planet formation. The goal of our work is to create an automated system that has the ability to efficiently process and classify large-scale astronomical datasets that missions such as WFIRST will produce. In this paper, we discuss our framework that utilizes features election and parameter optimization for classification models to automatically differentiate the different types of stellar variability and detect microlensing events.

Shvartzvald, Yossi

Temporal variability of the surface and atmosphere of Mars: Viking Orbiter color observations

We are near the final stages in the processing of a large Viking Orbiter global color dataset. Mosaics from 57 spacecraft revolutions (or 'revs' hereafter) were produced, most in both red and violet or red, green, and violet filters. Phase angles range from 13 deg to 85 deg. A total of approximately 2000 frames were processed through radiometric calibration, cosmetic cleanup, geometric control, reprojection, and mosaicking into single-rev mosaics at a scale of 1 km/pixel. All of the mosaics are geometrically tied to the 1/256 deg/pixel Mars Digital Image Mosaic (MDIM). Photometric normalization is in progress, to be followed by production of a 'best coverage' global mosaic at a scale of 1/64 deg/pixel (0.923 km/pixel). Global coverage is near 100 percent in red-filter mosaics and 98 percent and 60 percent in corresponding violet- and green-filter mosaics, respectively. Soon after completion, all final datasets (including single-rev mosaics) will be distributed to the planetary community on compact disks.

Mcewen, A. S.

Deep Learning System for Efficient Processing of Geostationary Satellite Imagery

Improved capabilities of Earth monitoring satellites are enabling a wide range of studies on the environmental effects of climate change, often leveraging the recent advancements in machine learning. At the same time, the new capabilities, including higher spatial resolution and temporal frequency, are expanding the amount of data generated at exponential rates. Further, a large majority of archived datasets generated by scientific processing is never used. This motivates the development of an efficient machine learning system for end-to-end processing of multi-level satellite datasets, from level 1 top of atmosphere observations to user friendly environmental variables of interest. Using current generation geostationary satellites GOES-16/17 (NOAA/NASA), and Himawari-8/9 (JAXA), we present an interchangeable set of machine models to perform spectral adjustment among sensors, physical model emulation, LEO-GEO emulation, and optical flow in a high performance computing environment. We use these tools on the NASA Earth eXchange (NEX) to generate consistent virtual observations across sensors, perform atmospheric correction and cloud detection, and estimate surface reflectance, surface temperature and atmospheric winds. This approach aims to improve the robustness of remotely sensed data processing by learning from diverse sets of observations while enabling near real-time and on-demand capabilities.

Thomas Vandal

ExaCA v2.0: A versatile, scalable, and performance portable cellular automata application for additive manufacturing solidification

The previously established ExaCA software for performance portable alloy grain structure simulation has been updated to better represent the solidification behavior during complex alloy processing conditions, such as those encountered during metal additive manufacturing (AM), and for improved performance and scalability. Here, an extension to the time–temperature history input data format and the core ExaCA algorithm to include an arbitrary number of melting and solidification events yielded improved prediction of texture for various melt pool geometries, expanding the range of AM-relevant conditions that can be accurately simulated. Improved heat transport process simulation coupling, including the creation of large raster datasets from single track time–temperature history data and in-memory coupling with the new, performance portable finite difference code Finch, were also demonstrated in example studies on the effect of multilayer AM microstructure predictions on hatch spacing and cell size, respectively. Additional new features are detailed and demonstrated, including the ability to perform simulations using various interfacial response function forms, execute simulations on state-of-the-art hardware, improved usability through post-processing versatility, and improved strong and weak scaling performance. The performance, physics, and versatility improvements demonstrated here will further enable large-scale studies on AM process–microstructure relationships that were not previously possible. Furthermore, the usability improvements and ability to run coupled AM process–microstructure simulations using the Finch-ExaCA workflow will facilitate broader use of this open-source software by the computational materials community.

36 MATERIALS SCIENCE

GeoNEX-ML: A Machine Learning System for Geostationary Satellite Imagery

Improved capabilities of earth monitoring satellites are enabling a wide range of studies on the environmental effects of climate change, often leveraging the recent advancements in machine learning. At the same time, the new capabilities, including higher spatial resolution and temporal frequency, are expanding the amount of data generated at exponential rates. Further, a large majority of archived datasets generated by scientific processing is never used. This motivates the development of an efficient machine learning system for end-to-end processing of multi-level satellite datasets, from level 1 top of atmosphere observations to user friendly environmental variables of interest. Using current generation geostationary satellites GOES-16/17 (NOAA/NASA), Himawari-8/9 (JAXA), and GK-2A (Korea), we present an interchangeable set of machine models to perform spectral adjustment, physical model emulation, LEO-GEO emulation, and optical flow in a high performance computing environment. We use these tools to generate consistent virtual observations across sensors, perform atmospheric correction and cloud detection, and estimate land surface temperature and atmospheric winds. This approach aims to improve the robustness of remotely sensed data processing by learning from diverse sets of observations while enabling near real-time and on-demand capabilities.

Geostationary satellites

TPCpp-10M: Simulated proton-proton collisions in a time projection chamber for AI foundation models

Scientific foundation models hold great promise for advancing nuclear and particle physics by improving analysis precision and accelerating discovery. Yet, progress in this field is often limited by the lack of openly available large scale datasets, as well as standardized evaluation tasks and metrics. Furthermore, the specialized knowledge and software typically required to process particle physics data pose significant barriers to interdisciplinary collaboration with the broader machine learning community. This work introduces a large, openly accessible dataset of 10 million simulated proton-proton collisions, designed to support self-supervised training of foundation models. To facilitate ease of use, the dataset is provided in a common NumPy format. In addition, it includes 70,000 labeled examples spanning three well defined downstream tasks: track finding, particle identification, and noise tagging, to enable systematic evaluation of the foundation model's adaptability. The simulated data are generated using the Pythia Monte Carlo event generator at a center of mass energy of $\sqrt{s}$ = 200 GeV and processed with Geant4 to include realistic detector conditions and signal emulation in the sPHENIX Time Projection Chamber at the Relativistic Heavy Ion Collider, located at Brookhaven National Laboratory. This dataset resource establishes a common ground for interdisciplinary research, enabling machine learning scientists and physicists alike to explore scaling behaviors, assess transferability, and accelerate progress toward foundation models in nuclear and high energy physics. The complete simulation and reconstruction chain is reproducible with the sPHENIX software stack. All data and code locations are provided under Data Accessibility.

Data Analysis, Statistics and Probability (physics

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Biochemistry & Molecular Biology

HPC Campaign Management: Remote data access with user-defined error bound using ADIOS and ZFP

Remote access to large-scale scientific datasets, like those generated by combustion simulations or other high-performance computing (HPC) applications, presents a significant challenge. Downloading entire datasets is often impractical due to their size and the bandwidth limitations of typical networks. To address this challenge, we propose a novel approach that enables efficient remote access to large datasets distributed across multiple facilities. Our method enables technologies to download only the data values of a select variable, in a select region of interest, to a user-defined accuracy. For this purpose, we extended the ADIOS IO library to provide read functions with user-defined accuracy, a remote data server that understands multidimensional selections of specific variables, steps and accuracy from an ADIOS dataset, and which uses lossy compression on the remote site to reduce the data to be transferred back to the client. In addition, our extension of the ADIOS library collects metadata from multiple datasets in small files called Campaign Archives, which can be shared among project participants on any HPC, cloud or laptop, and which can easily facilitate the discovery of content and pointers to the data location as well as remote access to the data by local tools as if data was local. This feature called Campaign Management, enables a group of scientists to manage related datasets stored in multiple files, across multiple facilities as if it was in a single file/database. We demonstrate the effectiveness of our approach using a 1.5 TB dataset from the S3D combustion simulation on Frontier at the Oak Ridge Leadership Facility. Even a single variable from this dataset, at 64 GB, is too large to be processed on a standard laptop. We show two different reading patterns for 2D plots and 3D visualization, with careful settings that a scientist studying combustion data would do and show that running the same Python scripts on Frontier directly takes comparable time than running them on the local laptop with remote access to the data on Frontier.

Podhorszki, Norbert [ORNL] (ORCID:000000019647542X

Modeling Global Biogenic Emission of Isoprene: Exploration of Model Drivers

Vegetation provides the major source of isoprene emission to the atmosphere. We present a modeling approach to estimate global biogenic isoprene emission. The isoprene flux model is linked to a process-based computer simulation model of biogenic trace-gas fluxes that operates on scales that link regional and global data sets and ecosystem nutrient transformations Isoprene emission estimates are determined from estimates of ecosystem specific biomass, emission factors, and algorithms based on light and temperature. Our approach differs from an existing modeling framework by including the process-based global model for terrestrial ecosystem production, satellite derived ecosystem classification, and isoprene emission measurements from a tropical deciduous forest. We explore the sensitivity of model estimates to input parameters. The resulting emission products from the global 1 degree x 1 degree coverage provided by the satellite datasets and the process model allow flux estimations across large spatial scales and enable direct linkage to atmospheric models of trace-gas transport and transformation.

Alexander, Susan E.

X-ray tomography of damage dynamics in advanced materials using a laser wakefield accelerator

Additively manufactured (AM) metals offer the potential for customizable, cost-effective components, but qualification and certification are crucial. Key to this process is understanding pore dynamics under stress, typically analyzed using micro-computed tomography. This study introduces laboratory-scale “betatron” x-rays from laser wakefield acceleration as a high-throughput alternative for x-ray tomography of advanced materials, such as AM AlSi10Mg alloys. Coupled with 3D finite element modeling, this method provides detailed insights into stress-porosity interactions. The approach delivers high-resolution scans, revealing that pore shape and local triaxiality significantly influence fracture dynamics, supporting advanced material characterization. This work also demonstrates the potential and versatility of laser-betatron x-ray μCT for generating large datasets to accelerate our understanding of the stochastic, process-specific nature of pore formation in AM alloys.

Senthilkumaran, Vigneshvar

In-Situ Process Monitoring, Synchronization, and Mapping Laser Powder Bed Fusion Builds of Ti6Al4V

The use of in-situ process monitoring is of interest to lower the cost of inspection for the qualification of laser powder bed fusion (LPBF) parts. Precise monitoring of the LPBF-AM build process constitutes a multi-scale and multi-discipline task. There are several significant challenges to the in-situ approach: the synchronization of sensor signals to process steps, the physical interpretation and classification of sensor signals, managing very large datasets, and comparing the inputs with the observed monitoring signals. At NASA Langley Research Center, a configurable architecture additive testbed has been developed to monitor the build process with synchronized sensors. The philosophy and method adopted for the synchronization of the cameras with laser power & position Ti-6Al-4V LPBF are described. The synchronized in-situ monitoring signals are compared with ex-situ nondestructive inspection and optical microscopy observations. Such comparisons permit a better understanding of how the sequential process actions of LPBF-AM can affect build quality.

Laser Powder Bed Fusion