Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Large Dataset Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Signal processing and spectral modeling for the BeEST experiment

The Beryllium Electron capture in Superconducting Tunnel junctions (BeEST) experiment searches for evidence of heavy neutrino mass eigenstates in the nuclear electron capture decay of 7 Be by precisely measuring the recoil energy of the 7 Li daughter. In Phase III, the BeEST experiment has been scaled from a singl superconducting tunnel junction (STJ) sensor to a 36-pixel array to increase sensitivity and mitigate gamma-induced backgrounds. Phase III also uses a new continuous data acquisition system that greatly increases the flexibility for signal processing and data cleaning. Here, we have developed procedures for signal processing and spectral fitting that are sufficiently robust to be automated for large datasets. Furthermore, this article presents the optimized procedures before unblinding the majority of the Phase III dataset to search for physics beyond the standard model.

6 ≤ A ≤ 19

OPEN-Augmented Reality GUI for Bioenergy Crop Phenotyping and Precision Agriculture (Donald Danforth Plant Science Center Final Scientific Technical Report)

The project led by the Donald Danforth Plant Science Center, in collaboration with Arizona State University, George Washington University, and Saint Louis University, has made significant strides in advancing the phenotypic analysis of bioenergy crops through the development of an innovative AI processing pipeline. This initiative was primarily funded by ARPA-E, with additional cost-sharing provided by the participating institutions. The project successfully utilized a variety of sensors—3D scanners, thermal, RGB, and hyperspectral—to refine algorithms for data-driven trait signature identification and improve the classification and visualization of plant traits. The developed AI processing pipeline is capable of handling the complex, multidimensional data characteristic of dynamic agricultural environments. 1) Contributions to understanding: The research has advanced the field of plant phenomics by showcasing the synergistic use of various sensor data to enhance the precision of trait analysis in bioenergy crops. Through the integration of 3D scanners, thermal, RGB, and hyperspectral sensors, the project has developed robust data-driven trait signature algorithms and visualization techniques. These innovations have facilitated detailed monitoring and management of plant traits, providing vital insights into plant growth dynamics and stress responses. Further, the project has broadened our understanding of how machine learning can be effectively applied in multi-sensor environments to refine trait analysis. By leveraging diverse datasets, the research has not only improved the accuracy of phenotypic assessments but also established a versatile methodological framework that can be extended beyond agriculture to other fields requiring detailed phenotypic analysis. 2) Technical effectiveness and economic feasibility: The AI processing pipeline developed in this project demonstrated significant technical effectiveness, achieving high throughput analysis of extensive phenotypic data and meeting targeted accuracies. This system exemplified the capability of advanced machine learning technologies to efficiently manage and analyze large, complex datasets. Economically, the implementation of the project-developed pipelines may offer substantial cost savings across multiple sectors. It enhances data analysis processes and significantly reduces the need for manual data interpretation, thereby decreasing both the time and resources required. 3) Public benefit: The project has significantly broadened the scope of agricultural methodologies to enhance phenotypic analysis, with potential applications in various sectors beyond agriculture. Additionally, the initiative fostered an enriching educational and collaborative environment, significantly enhancing the technical skills of participants. It also made substantial contributions to the scientific community by providing open-access data sets and tools, encouraging ongoing research and development across various disciplines. Overall, the project not only met its scientific goals but also showcased the extensive utility of integrating advanced machine learning and sensor data analysis technologies. These advancements have proven instrumental in driving forward both theoretical research and practical applications, setting a strong foundation for future explorations and innovations in data-driven science.

60 APPLIED LIFE SCIENCES

CryoDRGN-AI: neural ab initio reconstruction of challenging cryo-EM and cryo-ET datasets

Proteins and other biomolecules form dynamic macromolecular machines that are tightly orchestrated to move, bind, and perform chemistry. Cryo-electron microscopy (cryo-EM) and cryo-electron tomography (cryo-ET) can access the intrinsic heterogeneity of these complexes and are therefore key tools for understanding their function. However, 3D reconstruction of the collected imaging data presents a challenging computational problem, especially without any starting information, a setting termed ab initio reconstruction. Here, in this study, we introduce cryoDRGN-AI, a method leveraging an expressive neural representation and combining an exhaustive search strategy with gradient-based optimization to process challenging heterogeneous datasets. Using cryoDRGN-AI, we reveal new conformational states in large datasets, reconstruct previously unresolved motions from unfiltered datasets, and demonstrate ab initio reconstruction of biomolecular complexes from in situ data. With this expressive and scalable model for structure determination, we hope to unlock the full potential of cryo-EM and cryo-ET as a high-throughput tool for structural biology and discovery.

Levy, Axel [Stanford Univ., CA (United States); SL

Lithologic discrimination of volcanic and sedimentary rocks by spectral examination of Landsat TM data from the Puma, Central Andes Mountains

The Central Andes are widely used as a modern example of noncollisional mountain-building processes. The Puna is a high plateau in the Chilean and Argentine Central Andes extending southward from the altiplano of Bolivia and Peru. Young tectonic and volcanic features are well exposed on the surface of the arid Puna, making them prime targets for the application of high-resolution space imagery such as Shuttle Imaging Radar B and Landsat Thematic Mapper (TM). Two TM scene quadrants from this area are analyzed using interactive color image processing, examination, and automated classification algorithms. The large volumes of these high-resolution datasets require significantly different techniques than have been used previously for the interpretation of Landsat MSS data. Preliminary results include the determination of the radiance spectra of several volcanic and sedimentary rock units and the use of the spectra for automated classification. Structural interpretations have revealed several previously unknown folds in late Tertiary strata, and key zones have been targeted to be investigated in the field. The synoptic view of space imagery is already filling a critical gap between low-resolution geophysical data and traditional geologic field mapping in the reconnaissance study of poorly mapped mountain frontiers such as the Puna.

Fielding, E. J.

Comparison of automated chemical-guided segmentation and human annotation of soil organic matter in X-ray microcomputed tomography imaging in contrasted soil types

Soil organic matter (OM) formation and persistence is strongly influenced by the spatial distribution of organic substrates and microscale soil heterogeneity by dictating OM accessibility to microorganisms. However, traditional size and/or density fractionation techniques disrupt aggregate architecture, eliminating spatial information needed to fully understand intra-aggregate OM distribution. To quantify three-dimensional OM spatial distribution and automate segmentation in X-ray microcomputed tomography (µCT) imaging without human annotation bias, we developed an iodine gas vapor (I2) based staining workflow that eliminates labor-intensive manual annotation while maintaining segmentation accuracy, using aggregates from four taxonomically diverse soils (Xerofluvent, Haploxeroll Sphagnofibrist, Palehumult) with an 8-fold range of soil organic carbon. Human annotation of 10 µCT slices by the experienced and inexperienced annotators resulted in variations up to 3% in the Dice similarity coefficient (DSC), reflecting a degree of inherent subjectivity of manual labeling. Such inconsistencies are expected to compound as the number of manually annotated slices increases. Dual-energy µCT imaging at 33.1 keV (below the iodine (I) K-edge) and 33.2 keV (above the I K-edge) was used to resolve aggregate microstructure following I2 staining. The automated image subtraction pipeline identified OM regions by the I Kedge induced brightness increases, achieving DSC values of 0.58–0.83 relative to an experienced annotator. Sensitivity analyses revealed that the reconstruction alpha value—optimized via the open-source tool TomocuPy—and the 3D registration slice count were the primary determinants of accuracy, providing a novel benchmark for dual-energy soil imaging. The pipeline without GPU acceleration achieved 9.6 to 43.2 times faster than manual annotation. Using GPU-accelerated image post-processing and affine transformation matrices, the pipeline successfully segmented OM elements for large-scale datasets (3232×3232 pixel, 2048 slices) within ~5200 s from raw file acquisition to segmented output. The high-throughput approach enables the quantification of OM spatial distribution across diverse and heterogeneous soil.

Soil microbial biomass

Forecasting high-dimensional spatio-temporal systems from sparse measurements

This paper introduces a new neural network architecture designed to forecast high-dimensional spatio-temporal data using only sparse measurements. The architecture uses a two-stage end-to-end framework that combines neural ordinary differential equations (NODEs) with vision transformers. Initially, our approach models the underlying dynamics of complex systems within a low-dimensional space; and then it reconstructs the corresponding high-dimensional spatial fields. Many traditional methods involve decoding high-dimensional spatial fields before modeling the dynamics, while some other methods use an encoder to transition from high-dimensional observations to a latent space for dynamic modeling. In contrast, our approach directly uses sparse measurements to model the dynamics, bypassing the need for an encoder. This direct approach simplifies the modeling process, reduces computational complexity, and enhances the efficiency and scalability of the method for large datasets. We demonstrate the effectiveness of our framework through applications to various spatio-temporal systems, including fluid flows and global weather patterns. Although sparse measurements have limitations, our experiments reveal that they are sufficient to forecast system dynamics accurately over long time horizons. Our results also indicate that the performance of our proposed method remains robust across different sensor placement strategies, with further improvements as the number of sensors increases. This robustness underscores the flexibility of our architecture, particularly in real-world scenarios where sensor data is often sparse and unevenly distributed.

97 MATHEMATICS AND COMPUTING

A Downscaling Analysis of the Urban Influence on Rainfall: TRMM Satellite Component AMS Conference on Satellite Meteorology and Oceanography

A recent publication by Shepherd et al. (2002) demonstrated the feasibility of using TRMM precipitation radar (PR) estimates to identify precipitation anomalies caused by urbanization. The approach is particularly useful for investigating this global process because TRMM data span large portions of the globe and comprise an extended temporal dataset. Recent literature suggests that urbanized regions of Houston, Texas may be influencing lightning and precipitation formation over and downwind of the city. Possible mechanisms include: (1) enhanced convergence through interactions between the sea breeze, Galveston bay breeze, and urban heat island circulations, (2) enhanced convergence due to increased surface roughness over the city and/or destabilization of the boundary layer by the UHI, or (3) enhanced cloud condensation nuclei due to urban and industrial aerosol sources. In this study, a downscaling analysis of spatial and temporal trends in rainfall around the Houston Area is being conducted. The downscaling analysis concept involves identifying and quantifying urban rainfall anomalies at progressively smaller spatial and temporal scales using the TRMM satellite, ground-based radar, and a dense network of rain gauges. The goal is to test the hypothesis that the Houston urban district and regions in the climatological downwind region of the city exhibit enhanced rainfall amounts relative to the climatological upwind regions. TRMM was launched in 1997 and currently operates in a low-inclination (35 deg), non-sun-synchronous orbit at an altitude of 402 km (350 km prior to August 2001). The satellite analysis follows the methodologies described in Shepherd et al. (2002). Nearly five years of TRMM PR-derived mean monthly rainfall estimates are utilized to produce annual and warm season isohyetal analyses around Houston. Early results indicate that rainfall rates (mm/h) for the entire period are largest within 100 km northeast and east of Houston (e.g. the "hypothesized downwind region"). The mean rainfall rate over the Houston urban center is 30.5% larger than the upwind control region. The mean rainfall rate in the downwind region is 34.4% larger than the upwind region. An analysis of a parameter called the urban rainfall ratio (URR) illustrates that 65% (88%) of the satellite-derived rainfall rates in the downwind (upwind control) region are greater (less) than the mean background rainfall rate of the entire study region. When the data is stratified by summer months from 1998 to 2001 (June-August), even greater influence over and downwind of the urban area is observed in the statistics. This result is consistent with published reports of urban-generated rainfall being more prevalent in the warm season. The research demonstrates that the evolving TRMM satellite climatology is a credible way to detect mesoscale precipitation signatures that may be linked to urbanization. Early results also corroborate recent findings on Houston-induced convection/drainfall anomalies. Burian and Shepherd will report on other aspects of the downscaling analysis in future forums, but early rain gauge results are consistent with the satellite-based observations.

Shepherd, J. Marshall

Creating Benchmark Data for Artificial Intelligence and Machine Learning Space Biology Research

To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.

James Casaletto

Characterizing How Meteorological Forcing Selection and Parameter Uncertainty Influence Community Land Model Version 5 Hydrological Applications in the United States

Despite the increasing use of large-scale Land Surface Models (LSMs) in predicting hydrological responses in extreme conditions, there's a critical gap in understanding the uncertainties in these predictions. This study addresses this gap through a detailed diagnostic evaluation of the uncertainties arising from meteorological forcing selection and model parametrization in hydrological simulations of the Community Land Model version 5 (CLM5). CLM5 is configured at a spatial scale of about 12-km to simulate runoff processes for 464 headwater watersheds, selected from the Catchment Attributes for Large-Sample Studies (CAMELS) dataset to be representative of physiographic and climatic gradients across the conterminous United States. For each watershed, CLM5 is driven by five commonly used gridded forcing datasets in combination with a large ensemble (> 1200) of key CLM5 hydrologic parameters. Our results suggest that uncertainty in CLM5 runoff simulations resulting from both forcing and parametric sources is markedly higher in arid regions, e.g., Great Plains and Midwest regions. Uncertainty in low flow is dominated by parametric uncertainty, while the selection of meteorological forcing contributes more dominantly to high flow and seasonal flows during fall and spring. Our analysis also demonstrates that the selection of forcing datasets and the metrics used to calibrate CLM5 significantly impact the model’s predictive accuracy in extreme event severity for both floods and droughts. Overall, the results from this study highlight the need to understand and account for forcing and parametric uncertainties in CLM5 simulations, particularly for hazard and risk assessments addressing hydrologic extremes.

54 ENVIRONMENTAL SCIENCES

Data Mining for Science of the Sun-Earth Connection as a Single System

Establishing the Sun-Earth connection requires overcoming the challenges of exploring the data from past and current missions and leveraging tools and models (data mining) to create an efficient system treatment of the Sun and heliosphere. However, solar and heliospheric environment data constitute a vast source of information whose potential is far from being optimally exploited. In the next decade, the solar and heliospheric community will have to manage the increasing amount of information coming from new missions, improve reanalysis of data from past and current missions, and create new data products from the application of new methodologies. This complex task is further complicated by practical challenges such as different datasets and catalogs in different formats that may require different pre-processing and analysis tools, and the need for numerous analysis approaches that are not all fully optimized for large volumes of data. While several ongoing efforts aim at addressing these problems, the available datasets and tools are not always used to their full potential often due to lack of awareness of available resources. In this paper, we summarize the issues raised and goals discussed by members of the community during recent conference sessions focused on data mining for science.

Sun-Earth connection

Denoising Seismograms in the Time Domain Using a Deep Learning Model

Deep learning has emerged as a transformative tool for enhancing the extraction of reliable information from seismograms, addressing the increasing demand for precise and efficient seismic data analysis. We introduce an innovative encoder–decoder deep learning model, named WaveDenoiser, designed for noise reduction in the time domain, thereby eliminating the need for spectrogram computations that have been used for existing deep learning tools and significantly improving processing speed. Utilizing the benchmark dataset that is Stanford Earthquake Dataset, we developed three models of varying sizes: base, medium, and large. Notably, the large (referred to as WaveDenoiser) model demonstrated superior performance, achieving a median signal‐to‐noise ratio improvement of 8.8 dB on in‐distribution unseen data (in the same geographic region) and 7.7 dB on out‐distribution unseen data (in a new geographic region), outpacing both the base and medium models. Further evaluation of the WaveDenoiser model revealed a reduction in median arrival‐time errors by 0.02 s for P waves and 0.01 s for S waves when processing waveforms prior to phase picking using PhaseNet on in‐distribution unseen data. When tested on out‐distribution unseen data, the model also effectively reduced the P‐wave median arrival‐time error by 0.02 and 0.01 s in median arrival‐time error for S waves. Importantly, the application of WaveDenoiser resulted in a significant reduction of phase picking outliers by 1.1% to 3.6% for both P and S waves. In addition, we achieved over five times acceleration in processing speed compared with the seisBench implementation of DeepDenoiser. Our findings underscore the potential of WaveDenoiser as a powerful tool for improving seismic data analysis and processing efficiency.

P-waves

Avian Radar / Processed Data

This dataset contains radar track data of avian targets from DeTect's 7360 s-band radar during the large barge deployment from June 2024 though September 2024.

17 WIND ENERGY

MODIS Observations of Enhanced Clear Sky Reflectance Near Clouds

Several recent studies have found that the brightness of clear sky systematically increases near clouds. Understanding this increase is important both for a correct interpretation of observations and for improving our knowledge of aerosol-cloud interactions. However, while the studies suggested several processes to explain the increase, the significance of each process is yet to be determined. This study examines one of the suggested processes three-dimensional (3-D) radiative interactions between clouds and their surroundings by analyzing a large dataset of MODIS (Moderate Resolution Imaging Spectroradiometer) observations over the Northeast Atlantic Ocean. The results indicate that 3-D effects are responsible for a large portion of the observed increase, which extends to about 15 km away from clouds and is stronger (i) at shorter wavelengths (ii) near optically thicker clouds and (iii) near illuminated cloud sides. This implies that it is important to account for 3-D radiative effects in the interpretation of solar reflectance measurements over clear regions in the vicinity of clouds.

Varnai, Tamas

Climate Model Diagnostic Analyzer

The comprehensive and innovative evaluation of climate models with newly available global observations is critically needed for the improvement of climate model current-state representation and future-state predictability. A climate model diagnostic evaluation process requires physics-based multi-variable analyses that typically involve large-volume and heterogeneous datasets, making them both computation- and data-intensive. With an exploratory nature of climate data analyses and an explosive growth of datasets and service tools, scientists are struggling to keep track of their datasets, tools, and execution/study history, let alone sharing them with others. In response, we have developed a cloud-enabled, provenance-supported, web-service system called Climate Model Diagnostic Analyzer (CMDA). CMDA enables the physics-based, multivariable model performance evaluations and diagnoses through the comprehensive and synergistic use of multiple observational data, reanalysis data, and model outputs. At the same time, CMDA provides a crowd-sourcing space where scientists can organize their work efficiently and share their work with others. CMDA is empowered by many current state-of-the-art software packages in web service, provenance, and semantic search.

cloud computing

Geometry-complete diffusion for 3D molecule generation and optimization

Abstract Generative deep learning methods have recently been proposed for generating 3D molecules using equivariant graph neural networks (GNNs) within a denoising diffusion framework. However, such methods are unable to learn important geometric properties of 3D molecules, as they adopt molecule-agnostic and non-geometric GNNs as their 3D graph denoising networks, which notably hinders their ability to generate valid large 3D molecules. In this work, we address these gaps by introducing the Geometry-Complete Diffusion Model (GCDM) for 3D molecule generation, which outperforms existing 3D molecular diffusion models by significant margins across conditional and unconditional settings for the QM9 dataset and the larger GEOM-Drugs dataset, respectively. Importantly, we demonstrate that GCDM’s generative denoising process enables the model to generate a significant proportion of valid and energetically-stable large molecules at the scale of GEOM-Drugs, whereas previous methods fail to do so with the features they learn. Additionally, we show that extensions of GCDM can not only effectively design 3D molecules for specific protein pockets but can be repurposed to consistently optimize the geometry and chemical composition of existing 3D molecules for molecular stability and property specificity, demonstrating new versatility of molecular diffusion models. Code and data are freely available on GitHub .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Satellite orbit and data sampling requirements

Climate forcings and feedbacks vary over a wide range of time and space scales. The operation of non-linear feedbacks can couple variations at widely separated time and space scales and cause climatological phenomena to be intermittent. Consequently, monitoring of global, decadal changes in climate requires global observations that cover the whole range of space-time scales and are continuous over several decades. The sampling of smaller space-time scales must have sufficient statistical accuracy to measure the small changes in the forcings and feedbacks anticipated in the next few decades, while continuity of measurements is crucial for unambiguous interpretation of climate change. Shorter records of monthly and regional (500-1000 km) measurements with similar accuracies can also provide valuable information about climate processes, when 'natural experiments' such as large volcanic eruptions or El Ninos occur. In this section existing satellite datasets and climate model simulations are used to test the satellite orbits and sampling required to achieve accurate measurements of changes in forcings and feedbacks at monthly frequency and 1000 km (regional) scale.

Rossow, William

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING

SNAPRed: Reduction of multidimensional neutron time-of-flight diffraction data

SNAP is a neutron time-of-flight diffractometer at the Spallation Neutron Source operated by Oak Ridge National Laboratory. It generates large arrays of neutron detection events that encode the crystalline atomic structure of materials under study. SNAPRed is an application that makes these datasets accessible to end users by orchestrating the process of data reduction while automatically managing the variable neutron instrumentation configuration. It supports arbitrary grouping and masking of individual detector pixels and includes custom-developed data compression approaches to accommodate the large volumes of data generated by the SNAP instrument.

Diffraction