Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Science Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Legacy Survey of Space and Time Data Preview 2: source dataset type

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of the source dataset type. These are measurements for detected sources in processed visit images. This release contains 28,589 datasets of this type.

79 ASTRONOMY AND ASTROPHYSICS

Generalization error guaranteed auto-encoder-based nonlinear model reduction for operator learning

Many physical processes in science and engineering are naturally represented by operators between infinite-dimensional function spaces. The problem of operator learning, in this context, seeks to extract these physical processes from empirical data, which is challenging due to the infinite or high dimensionality of data. An integral component in addressing this challenge is model reduction, which reduces both the data dimensionality and problem size. In this paper, we utilize low-dimensional nonlinear structures in model reduction by investigating Auto-Encoder-based Neural Network (AENet). AENet first learns the latent variables of the input data and then learns the transformation from these latent variables to corresponding output data. Our numerical experiments validate the ability of AENet to accurately learn the solution operator of nonlinear partial differential equations. Furthermore, we establish a mathematical and statistical estimation theory that analyzes the generalization error of AENet. Finally, our theoretical framework shows that the sample complexity of training AENet is intricately tied to the intrinsic dimension of the modeled process, while also demonstrating the robustness of AENet to noise.

Auto-encoder

Stream discharge and temperature data collected within the East and Taylor Watershed, Colorado for the Lawrence Berkeley National Laboratory Watershed Function Science Focus Area (water years 2019 to 2025)

This dataset contains stream discharge and temperature data for water years 2019 to 2025 from the East and Taylor Watersheds in Colorado, United States. This data was collected to understand hydrological processes occurring in the East River and Taylor River Watersheds, Colorado, which is part of the Lawrence Berkeley National Laboratory Watershed Function Scientific Focus Area. Data includes instantaneous observed discharge using salt dilution and acoustic doppler velocimeter techniques, raw pressure transducer downloaded data, sub-hourly temperature as well as corrected water level and associated stream discharge and mean daily values. Notes on water level corrections, rating curve development and metadata provided. A rating curve is the translation of depth to streamflow. The rating curve can be used as a quantitative measure of the “quality of the data.” Data within this dataset is formatted using ESS-DIVE’s Hydrological Monitoring Reporting Format. This data package contains (1) a zip file (Stream_Discharge_Data_WY19-WY25.zip) containing stream discharge and temperature data organized by location; (2) an InstallationMethods file (InstallationMethods.csv) describing metadata about the installation; (3) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; (4) a data dictionary (dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (5) a locations metadata file (locations.csv); (6) and a sensor metadata file (sensors.csv). All data files are in non-proprietary formats (csv, png, or pdf formats). Please contact Rosemary Carroll, Curtis Beutler, or Austin Shirley for any support in accessing the files. Update on 2023-05-12: Additional data from WYs 2021 and 2022 were added. Additionally, the dataset was converted using ESS-DIVE’s Hydrological Monitoring Reporting Format. Data files were reformatted to match reporting format guidance, new metadata files were added, and files were converted from excel to CSV. Update on 2025-05-16: Additional data from WYs 2022 (for locations not previously included), 2023, and 2024 were added. An additional descriptive PDF (WFSFA_Streamflow_Hydrograph_Disclaimer.pdf) was added. Metadata files were updated to reflect the addition of new data and locations. Update on 2026-05-18: Additional data from WY 2025 were added, including a new location Upper Trail Creek (TR-TCG2). Metadata files were updated to reflect the addition of new data.

54 ENVIRONMENTAL SCIENCES

Feature Based Qualification (FBQ) of Wire Arc Additively Manufactured (WAAM) 17-4PH Martensitic Stainless Steels

The Department of Defense (DOD) programs of records desire to reduce the time and cost of the development and delivery loop in metal additive manufacturing (AM), including establishing forwarded AM capabilities. The success of these efforts relies on a robust and qualified process. To achieve this, the United States Army Combat Capabilities Development Command Ground Vehicle Systems Center Materials Engineering (GVME) needs to be able to quickly evaluate, test, and develop feedstocks, processes, and parts. This report is directed towards demonstrating the need for a framework for metal AM processes, defining and exploring geometries for metal AM process qualification, testing resultant deposition, and delivering actionable data.

36 MATERIALS SCIENCE

Human Host Cellular Response to HCoV-229E Infection Transcriptomics (ACS-DP1)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5), immortalized human lung fibroblasts cells (MRC5) (MOI5), and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue. Sample data was acquired using an Illumina HiSeq 2000 sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES

Legacy Survey of Space and Time Data Preview 2: Source searchable catalog

We present Rubin Data Preview 2 (DP2), the second data preview from the NDF-DOE Vera C. Rubin Observatory. Data Preview 2 (DP2) comprises coadds, detection catalogs, and ancillary data products; and when fully released will also include single-epoch images and difference images. DP2 is derived from observations acquired by the LSST Science Camera (LSSTCam) on the Simonyi Survey Telescope at the Summit Facility on Cerro Pachón, Chile, primarily during the on-sky commissioning campaign between 2025-04-16 and 2025-09-21, supplemented by observations taken between 2025-10-25 and 2026-01-06 that overlap the commissioning footprint. The DP2 footprint comprises the Science Validation wide-area survey, five Deep Drilling Fields, and a number of targeted small-field regions, including Trifid and Lagoon, Prawn, M49, and New Horizons, all observed as part of the Rubin First Look campaign. Each field was imaged in up to six broad photometric bands, ugrizy, and coadded to produce deep imaging covering an estimated 3,000 deg2. The addition of single-visit-only areas expands the total DP2 footprint to an estimated 15,000 deg2, with coverage in at least one filter. The median per-visit PSF FWHM across the wide-area survey ranges from 1.17 arcsec in the z band to 1.26 arcsec in g and r bands. The deepest field, reaches estimated coadded 5σ depths of u=26 mag, g=26.8 mag, r=26.3 mag, i=26.1 mag, z=25.3 mag, y=23.9 mag. Based on a roughly five-month primary observing baseline and covering only part of the eventual LSST footprint, DP2's area, depth, and multiband coverage nonetheless support a broad range of early science investigations ahead of LSST Data Release This dataset is a subset of the full data release consisting of a searchable catalog named Source. This catalog contains measurements for detected sources in processed visit images. This catalog contains 17,566,180,086 rows with 154 columns.

79 ASTRONOMY AND ASTROPHYSICS

CROCUS Weather Data at Argonne National Laboratory Prairie Site

Vaisala WXT sensor is an all-in-one weather instrument that provides 6 of the most important weather parameters: barometric pressure, temperature, relative humidity, rainfall, wind speed and direction. Temperature, pressure, relative humidity, and rainfall are sampled at 1 second frequency, while wind speed/direction is measured at ten per second (10Hz) frequency. These measurements are useful for looking at characterizing local weather, identifying unique weather events, and studying local turbulence, especially given the high temporal resolution of the wind measurements. These measurements are collected at the Argonne Testbed for Multiscale Observational Science (ATMOS), a prairie field site at Argonne National Laboratory in Lemont, Illinois. Data is available in the netCDF data format, we encourage data users review documentation through Project Pythia to understand how to work with netCDF data https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html. The data is aggregated into daily frequency to make it easier to process multiple days, and compress the higher-resolution fields. Each file contains one day's worth of data (24 hours, starting at 0000 UTC). File naming convention includes the project (CROCUS), location (atmos), data level (raw, a1), date (year, month, day), and hour (0000).

54 ENVIRONMENTAL SCIENCES

BLOC Site - ASSIST Thermodynamic Retrievals TROPoe v0.18 / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). This is a post-processed dataset and recommended for use. The profiles are retrieved every 10 minutes from instantaneous radiances observed with an Atmospheric Sounder Spectrometer by Infrared Spectral Technology (ASSIST, Michaud-Belleau et al. 2025) operated by NOAA Physical Sciences Laboratory (PSL) on Block Island for WFIP3. The spectral bands used in the retrieval are in the wavenumber range from 612 - 905.4 cm-1 and are specified in Turner and Löhnert (2021). Additional input data in TROPoe are cloud base height from a collocated ceilometer operated by NOAA GML and temperature, water vapor mixing ratio, and pressure from a collocated surface tower operated by NOAA PSL. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Upton, NY. The TROPoe docker container (version 0.18) is available from Docker Hub at https://hub.docker.com/r/davidturner53/tropoe/tags, and the source code code is available in the GitHub repository https://github.com/OAR-atmospheric-observations/TROPoe.

17 WIND ENERGY

NANT Site - ASSIST Thermodynamic Retrievals TROPoe v0.18 / Derived Data

This dataset contains daily files with thermodynamic profiles retrieved with the optimal estimation physical retrieval TROPoe (Turner and Löhnert 2014; Turner and Blumberg 2019; Turner and Löhnert 2021). This is a post-processed dataset and recommended for use. The profiles are retrieved every 10 minutes from instantaneous radiances observed with an Atmospheric Sounder Spectrometer by Infrared Spectral Technology (ASSIST, Michaud-Belleau et al. 2025) operated by NOAA Physical Sciences Laboratory (PSL) on Nantucket Island for WFIP3. The spectral bands used in the retrieval are in the wavenumber range from 612 - 905.4 cm-1 and are specified in Turner and Löhnert (2021). Additional input data in TROPoe are cloud base height from a collocated ceilometer operated by NOAA GML and temperature, water vapor mixing ratio, and pressure from a collocated surface tower operated by NOAA PSL. In addition to these temporally resolved input data, TROPoe requires an a priori dataset (prior) that provides mean climatological estimates of thermodynamic profiles and specifies how temperature and humidity covary with height as an input (for details see, e.g., Djalalova et al. 2022). The prior is a key component of the retrieval and provides a constraint on the ill-posed inversion problem. A monthly prior was computed from operational radiosonde launches at Upton, NY. The TROPoe docker container (version 0.18) is available from Docker Hub at https://hub.docker.com/r/davidturner53/tropoe/tags, and the source code code is available in the GitHub repository https://github.com/OAR-atmospheric-observations/TROPoe.

17 WIND ENERGY

High-count rate effects in event processing for XRISM/ Resolve X-ray microcalorimeter: I. Ground test

The spectroscopic performance of an X-ray microcalorimeter is compromised at high count rates. We utilize the Resolve X-ray microcalorimeter onboard the XRISM satellite to examine the effects observed during high-count rate measurements and propose modeling approaches to mitigate them. We specifically address the following instrumental effects that impact performance: CPU limit, pile-up, and untriggered electrical cross-talk. Experimental data at high count rates were acquired during ground testing using the flight model instrument and a calibration X-ray source. In the experiment, data processing not limited by the performance of the onboard CPU was run in parallel, which cannot be done in orbit. This makes it possible to access the data degradation caused by limited CPU performance. We use these data to develop models that allow for a more accurate estimation of the aforementioned effects. To illustrate the application of these models in observation planning, we present a simulated observation of GX 13+1. Understanding and addressing these issues is crucial to enhancing the reliability and precision of X-ray spectroscopy in situations characterized by elevated count rates.

47 OTHER INSTRUMENTATION

A novel method may reveal bulk metallic glass compressive ductility trends in high data rate nanoindentation

Recent methods allow novel amorphous alloy compositions to be rapidly manufactured at small scale; however, obtaining materials properties such as compressive ductility from these smaller specimens has remained a challenge. Here, we suggest a potential high-throughput nanoindentation method that may be able to rapidly characterize the relative compressive ductility between these alloys based on their serration characteristics. The properties of emergent serrations, when interpreted in a simple micromechanical stress relaxation model, may order these materials by their compressive plastic strain to failure. These results are consistent with the ordering obtained from compressed specimens as well as with model simulations, suggesting that this model may be broadly useful for interpreting compressive ductility from nanoindentation serrations. After it is validated on more materials, this new method will match the rapid pace of amorphous alloy development, thus allowing metallic glass properties to be fine-tuned for each application prior to scale prototyping.

36 MATERIALS SCIENCE

3D Continuous Forcing Dataset from 3D Constrained Variational Analysis at SGP

The continuous 3D large-scale forcing (VARANAL3D) data set derived from 3D constrained variational analysis (3DCVA) extends the conventional constrained variational analysis method by incorporating multiple sub-columns within the analysis domain. This advancement introduces spatial variability into the large-scale forcing fields, thereby enriching the data set’s applicability. The VARANAL3D data set spans from 2004 to 2018 and covers a region of 5˚×4.5˚ domain around the ARM SGP site. The analysis domain is divided into 10×9 sub-columns with 0.5˚ resolution. The 3D large-scale forcing data provides necessary variables to drive and evaluate single-column models (SCM), cloud-resolving models (CRM) ,and large-eddy simulations (LES), as well as information for testing model sensitivity to spatial variability of the large-scale forcing data, facilitating more rigorous testing and refinement of physical processes in SCM/CRM/LES.

54 ENVIRONMENTAL SCIENCES

ZMPY3D: accelerating protein structure volume analysis through vectorized 3D Zernike moments and Python-based GPU integration

Abstract Motivation Volumetric 3D object analyses are being applied in research fields such as structural bioinformatics, biophysics, and structural biology, with potential integration of artificial intelligence/machine learning (AI/ML) techniques. One such method, 3D Zernike moments, has proven valuable in analyzing protein structures (e.g., protein fold classification, protein–protein interaction analysis, and molecular dynamics simulations). Their compactness and efficiency make them amenable to large-scale analyses. Established methods for deriving 3D Zernike moments, however, can be inefficient, particularly when higher order terms are required, hindering broader applications. As the volume of experimental and computationally-predicted protein structure information continues to increase, structural biology has become a “big data” science requiring more efficient analysis tools. Results This application note presents a Python-based software package, ZMPY3D, to accelerate computation of 3D Zernike moments by vectorizing the mathematical formulae and using graphical processing units (GPUs). The package offers popular GPU-supported libraries such as CuPy and TensorFlow together with NumPy implementations, aiming to improve computational efficiency, adaptability, and flexibility in future algorithm development. The ZMPY3D package can be installed via PyPI, and the source code is available from GitHub. Volumetric-based protein 3D structural similarity scores and transform matrix of superposition functionalities have both been implemented, creating a powerful computational tool that will allow the research community to amalgamate 3D Zernike moments with existing AI/ML tools, to advance research and education in protein structure bioinformatics. Availability and implementation ZMPY3D, implemented in Python, is available on GitHub (https://github.com/tawssie/ZMPY3D) and PyPI, released under the GPL License.

Lai, Jhih-Siang (ORCID:0000000156775890)

Chlamydomonas reinhardtii responses to Fe-excess, Fe-deficiency, and Fe-limitation in either photoautotrophic or mixotrophic growth

A systems level analysis of Chlamydomonas reinhardtii grown photoautotrophically or mixotrophically with a reduced carbon source, acetate, under four different defined Fe stages of Fe-replete, Fe-deficient, Fe-limited, or Fe-excess. Samples were digested with trypsin, labeled with TMT 10-Plex, then analyzed by LC-MS/MS. Data was searched with MS-GF+ using PNNL's DMS Processing pipeline. [doi:10.25345/C5707X12X] [dataset license: CC0 1.0 Universal (CC0 1.0)]

59 BASIC BIOLOGICAL SCIENCES

Human Primary Airway Epithelium +/- Macrophages Response to HCoV-229E Infection Transcriptomics (ACS-DP3)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-299E) infection. Sample data was obtained for mock and infected (MOI 3) primary human airway epithelial cells with and without macrophages and grown in air-liquid interface conditions. Sample data was acquired using an Illumina Hi-Seq 4000 sequencer system and further processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES

CROCUS Air Quality Data at Argonne National Laboratory Prairie Site

The AQT (Vaisala AQT530) instrument provides observations on meteorological conditions, including particulate matter (PM2.5, PM10), gas species concentrations (NO, NO2, O3, CO), and environment temperature and moisture. These measurements are critical for understanding air quality. These measurements are useful for understanding changes in aerosol properties, air quality research, and comparing to model experiments especially in urban environments. These measurements are collected at the Argonne Testbed for Multiscale Observational Science (ATMOS), a prairie field site at Argonne National Laboratory in Lemont, Illinois. Data is available in the netCDF data format, we encourage data users review documentation through Project Pythia to understand how to work with netCDF data https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html. Each file contains one day's worth of data (24 hours, starting at 0000 UTC). The data is aggregated into daily frequency to make it easier to process multiple days, and compress the higher-resolution fields. File naming convention includes the project (CROCUS), location (atmos), data level (raw, a1), date (year, month, day), and hour (0000).

54 ENVIRONMENTAL SCIENCES