Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessed data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Automatic Computer Mapping of Terrain

Computer processing of 17 wavelength bands of visible, reflective infrared, and thermal infrared scanner spectrometer data, and of three wavelength bands derived from color aerial film has resulted in successful automatic computer mapping of eight or more terrain classes in a Yellowstone National Park test site. The tests involved: (1) supervised and non-supervised computer programs; (2) special preprocessing of the scanner data to reduce computer processing time and cost, and improve the accuracy; and (3) studies of the effectiveness of the proposed Earth Resources Technology Satellite (ERTS) data channels in the automatic mapping of the same terrain, based on simulations, using the same set of scanner data. The following terrain classes have been mapped with greater than 80 percent accuracy in a 12-square-mile area with 1,800 feet of relief; (1) bedrock exposures, (2) vegetated rock rubble, (3) talus, (4) glacial kame meadow, (5) glacial till meadow, (6) forest, (7) bog, and (8) water. In addition, shadows of clouds and cliffs are depicted, but were greatly reduced by using preprocessing techniques.

Smedes, H. W.↗

Brady Geodatabase for Geothermal Exploration Artificial Intelligence

These files contain the geodatabases related to Brady's Geothermal Field. It includes all input and output files for the Geothermal Exploration Artificial Intelligence. Input and output files are sorted into three categories: raw data, pre-processed data, and analysis (post-processed data). In each of these categories there are six additional types of raster catalogs which are titled Radar, SWIR, Thermal, Geophysics, Geology, and Wells. These inputs and outputs were used with the Geothermal Exploration Artificial Intelligence to identify indicators of blind geothermal systems at the Brady Hot Springs Geothermal Site. The included zip file is a geodatabase to be used with ArcGIS and the tar file is an inclusive database that encompasses the inputs and outputs for the Brady Hot Springs Geothermal Site.

15 GEOTHERMAL ENERGY↗

Survey of Time Shift Detection Algorithms for Measured PV Data

In this research, three variations of time shift detection algorithms were tested for their ability to detect time shift issues (including daylight savings time and random time shifts) in measured PV data sets. Two algorithms from the Python PVAnalytics package were assessed, and one algorithm from the Solar-Data-Tools package was assessed. Each algorithm's ability to accurately detect and measure time shifts was assessed.

automated preprocessing↗

A Data-Driven Framework for Predicting the Sorting and Screening Performance of an Integrated Biomass Feedstock Preprocessing System

The characteristics of mechanically sorted and screened lignocellulosic biomass, such as the mass contents of corn stover anatomical fractions (leaves, husks, stalks, cobs, etc.), can be used to calculate the intermediate feedstock quality attributes “yield” and “purity” that indicate the conversion efficiency of biocrude. No prior study has investigated the correlations from the characteristics of raw biomass and preprocessing unit operation parameters to those intermediate feedstock quality attributes. This work presents a data-driven framework for assessing and predicting the intermediate feedstock quality attributes in an integrated biomass feedstock preprocessing system. Our study used corn stover as a typical type of herbaceous biomass because of its abundance in the U.S. It began with data acquisition of moisture content, particle size distribution, and anatomical fractions of the materials after each unit operation in the system. The objective of this preprocessing system is to minimize husks and leaves and maximizing cobs and stalks by mechanically separating the materials into three streams via disc screen and air separator. Prototype neural network models were then developed to evaluate the feasibility of predicting process outcomes based on measurable parameters. It is found that incorporating physical constraints into these prediction models significantly enhances the accuracy of the predicted yield and purity against the ground truth data. The experimental data and model predictions indicate that decreasing throughput increases purity, while higher throughput results in lower purity. Finally, an optimization problem was introduced to search optimal combinations of feed material properties and preprocessing unit operation parameters, as the intermediate feedstock quality attributes – yield and purity, appeared to be competing factors. The study also suggests the continual need to improve the data-driven framework’s predictability by incorporating more accurate physical models to describe the dynamics in the preprocessing units such as the air separator.

09 - BIOMASS FUELS↗

Day-Ahead Probabilistic Forecasting of Net-Load and Demand Response Potentials with High Penetration of Behind-the-Meter Solar-plus-Storage

The goal of this project is to develop advanced methods for day-ahead net-load forecasting, by leveraging the state-of-the-art machine learning techniques. The developed models produce both point and probabilistic forecasts for a variety of use cases, and are versatile to work with different types of data sets. The innovation lies in the novel design of the architectures, leveraging the most recent advances in machine learning that have not been explored in power systems, accompanied by techniques in the broader artificial intelligence fields such as fuzzy systems. This project has achieved the following accomplishments: (1) preprocessing of over 10 data sets covering varying geographical regions, time horizons, and system levels, which form a robust foundation for training and evaluating forecasting models across a wide range of realistic grid scenarios; (2) development of an interactive web app that enables exploratory analysis of load and generation data, and supports better understanding of data trends, anomalies, and correlations, facilitating model development and stakeholder engagement; (3) implementation of over 10 benchmark models for point and probabilistic forecasting, which include a mix of conventional machine learning methods and state-of-the-art deep learning approaches, providing a comprehensive baseline for performance comparison and validation of the proposed models; (4) development of a fuzzy system based gradient boosting model, tailored for small (less than 3 years) data sets, which achieves a mean absolute percentage error (MAPE) of 4% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (5) development of a Transformer (a state-of-the-art deep learning architecture) based neural network model, tailored for large (3 years or more) data sets, which achieves a MAPE of 2% for point forecasting and a 20% improvement in average pinball loss for probabilistic forecasting; (6) development of a methodology for quantifying DR potential, and extensions of the previous models for multi-target forecasting of net load and DR potential, which achieve a MAPE of 10% for DR potential.

24 POWER TRANSMISSION AND DISTRIBUTION↗

JOVE NASA-FIT program: Microgravity and aeronomy projects

This semi-annual status report is divided into two sections: Scanning Tunneling Microscopy Lab and Aeronomy Lab. The Scanning Tunneling Microscopy (STM) research involves studying solar cell materials using the STM built at Florida Tech using a portion of our initial Jove equipment funding. One result of the participation in the FSEC project will be to design and build an STM system which is portable. This could serve as a prototype STM system which might be used on the Space Shuttle during a Spacelab mission, or onboard the proposed Space Station. The scanning tunneling microscope is only able to image the surface structure of electrically conductive crystals; by building an atomic force microscope (AFM) the surface structure of any sample, regardless of its conductivity, will be able to be imaged. With regards to the Aeronomy Lab, a total of four different mesospheric oxygen emission codes were created to calculate the intensity along the line of sight of the shuttle observations for 2972A, Herzberg I, Herzberg II, and Chamberlain bands. The thermosphere-ionosphere coupling project was completed with two major accomplishments: collection of 500 data points on modulation of neutral wind with geophysical variables, and establishment of constraints on behavior of the height of the ionosphere as a result of interaction between geophysical and geometrical factors. The magnetotail plasma project has been centered around familiarization with the subject in the form of a literature search and preprocessing of IMP-8 data.

Patterson, James D.↗

I-GCN: A Graph Convolutional Network Accelerator with Runtime Locality Enhancement through Islandization

In this paper, we propose a novel hardware accelerator for GCN inference called I-GCN that significantly improves data locality and reduces unnecessary computation through a new online graph restructuring algorithm we refer to as islandization. The proposed algorithm finds clusters of nodes with strong internal but weak external connections. The islandization process yields two major benefits. First, by processing islands rather than individual nodes, there is better on-chip data reuse and fewer off-chip memory accesses. Second, there is less redundant computation as aggregation for common/shared neighbors in an island can be reused. The parallel search, identification, and leverage of graph islands are all handled purely in hardware at runtime working in an incremental pipelined manner. This is done without any preprocessing of the graph data or adjustment of the GCN model structure.

Geng, Tong↗

Data Arrays for Microearthquake (MEQ) Monitoring using Deep Learning for the Newberry EGS Sites

The 'Machine Learning Approaches to Predicting Induced Seismicity and Imaging Geothermal Reservoir Properties' project looks to apply machine learning (ML) methods to Microearthquake (MEQ) data for imaging geothermal reservoir properties and forecasting seismic events, in order to advance geothermal exploration and safe geothermal energy production. As part of the project, this submission provides data arrays for 149 microearthquakes between the year 2012 and 2013 at the Newberry EGS Site for use with the Deep Learning Algorithm that has been developed. The data provided includes raw waveform data, location data, normalized waveform data, and processed waveform data. Penn State Geothermal Team has shared the following files from the project: - 149 microearthquakes (MEQs) between 2012 and 2013 at Newberry EGS sites, 'Normalized Waveform Inputs.npz' are normalized waveforms. - labels of 149 MEQs: Processed Waveform Inputs.npz - location labels of 149 MEQs: Location Data.npz Note: .npz is the python file format by NumPy that provides storage of array data.

15 GEOTHERMAL ENERGY↗

In situ melt pool measurements for laser powder bed fusion using multi sensing and correlation analysis

Laser powder bed fusion is a promising technology for local deposition and microstructure control, but it suffers from defects such as delamination and porosity due to the lack of understanding of melt pool dynamics. To study the fundamental behavior of the melt pool, both geometric and thermal sensing with high spatial and temporal resolutions are necessary. This work applies and integrates three advanced sensing technologies: synchrotron X-ray imaging, high-speed IR camera, and high-spatial-resolution IR camera to characterize the evolution of the melt pool shape, keyhole, vapor plume, and thermal evolution in Ti–6Al–4V and 410 stainless steel spot melt cases. Aside from presenting the sensing capability, this paper develops an effective algorithm for high-speed X-ray imaging data to identify melt pool geometries accurately. Preprocessing methods are also implemented for the IR data to estimate the emissivity value and extrapolate the saturated pixels. Quantifications on boundary velocities, melt pool dimensions, thermal gradients, and cooling rates are performed, enabling future comprehensive melt pool dynamics and microstructure analysis. The study discovers a strong correlation between the thermal and X-ray data, demonstrating the feasibility of using relatively cheap IR cameras to predict features that currently can only be captured using costly synchrotron X-ray imaging. Such correlation can be used for future thermal-based melt pool control and model validation.

47 OTHER INSTRUMENTATION↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Concepts for on-board satellite image registration, volume 1

The NASA-NEEDS program goals present a requirement for on-board signal processing to achieve user-compatible, information-adaptive data acquisition. One very specific area of interest is the preprocessing required to register imaging sensor data which have been distorted by anomalies in subsatellite-point position and/or attitude control. The concepts and considerations involved in using state-of-the-art positioning systems such as the Global Positioning System (GPS) in concert with state-of-the-art attitude stabilization and/or determination systems to provide the required registration accuracy are discussed with emphasis on assessing the accuracy to which a given image picture element can be located and identified, determining those algorithms required to augment the registration procedure and evaluating the technology impact on performing these procedures on-board the satellite.

Ruedger, W. H.↗

LANDSAT 4 band 6 data evaluation

The objectives of this investigation are to evaluate and monitor the radiometric integrity of the LANDSAT-D Thematic Mapper (TM) thermal infrared channel (Band 6) data to develop improved radiometric preprocessing calibration techniques for removal of atmospheric effects. Efforts this period have concentrated on underflight data collection. Two successful flights were made on September 18 and October 6. The radiosonde data for these flights have been obtained.

Source record↗

Some practical universal noiseless coding techniques, part 3, module PSl14,K+

The algorithmic definitions, performance characterizations, and application notes for a high-performance adaptive noiseless coding module are provided. Subsets of these algorithms are currently under development in custom very large scale integration (VLSI) at three NASA centers. The generality of coding algorithms recently reported is extended. The module incorporates a powerful adaptive noiseless coder for Standard Data Sources (i.e., sources whose symbols can be represented by uncorrelated non-negative integers, where smaller integers are more likely than the larger ones). Coders can be specified to provide performance close to the data entropy over any desired dynamic range (of entropy) above 0.75 bit/sample. This is accomplished by adaptively choosing the best of many efficient variable-length coding options to use on each short block of data (e.g., 16 samples) All code options used for entropies above 1.5 bits/sample are 'Huffman Equivalent', but they require no table lookups to implement. The coding can be performed directly on data that have been preprocessed to exhibit the characteristics of a standard source. Alternatively, a built-in predictive preprocessor can be used where applicable. This built-in preprocessor includes the familiar 1-D predictor followed by a function that maps the prediction error sequences into the desired standard form. Additionally, an external prediction can be substituted if desired. A broad range of issues dealing with the interface between the coding module and the data systems it might serve are further addressed. These issues include: multidimensional prediction, archival access, sensor noise, rate control, code rate improvements outside the module, and the optimality of certain internal code options.

Rice, Robert F.↗

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Visualization techniques to aid in the analysis of multi-spectral astrophysical data sets

The goal of this project was to support the scientific analysis of multi-spectral astrophysical data by means of scientific visualization. Scientific visualization offers its greatest value if it is not used as a method separate or alternative to other data analysis methods but rather in addition to these methods. Together with quantitative analysis of data, such as offered by statistical analysis, image or signal processing, visualization attempts to explore all information inherent in astrophysical data in the most effective way. Data visualization is one aspect of data analysis. Our taxonomy as developed in Section 2 includes identification and access to existing information, preprocessing and quantitative analysis of data, visual representation and the user interface as major components to the software environment of astrophysical data analysis. In pursuing our goal to provide methods and tools for scientific visualization of multi-spectral astrophysical data, we therefore looked at scientific data analysis as one whole process, adding visualization tools to an already existing environment and integrating the various components that define a scientific data analysis environment. As long as the software development process of each component is separate from all other components, users of data analysis software are constantly interrupted in their scientific work in order to convert from one data format to another, or to move from one storage medium to another, or to switch from one user interface to another. We also took an in-depth look at scientific visualization and its underlying concepts, current visualization systems, their contributions, and their shortcomings. The role of data visualization is to stimulate mental processes different from quantitative data analysis, such as the perception of spatial relationships or the discovery of patterns or anomalies while browsing through large data sets. Visualization often leads to an intuitive understanding of the meaning of data values and their relationships by sacrificing accuracy in interpreting the data values. In order to be accurate in the interpretation, data values need to be measured, computed on, and compared to theoretical or empirical models (quantitative analysis). If visualization software hampers quantitative analysis (which happens with some commercial visualization products), its use is greatly diminished for astrophysical data analysis. The software system STAR (Scientific Toolkit for Astrophysical Research) was developed as a prototype during the course of the project to better understand the pragmatic concerns raised in the project. STAR led to a better understanding on the importance of collaboration between astrophysicists and computer scientists.

Brugel, Edward W.↗

Visualization techniques to aid in the analysis of multispectral astrophysical data sets

The goal of this project was to support the scientific analysis of multi-spectral astrophysical data by means of scientific visualization. Scientific visualization offers its greatest value if it is not used as a method separate or alternative to other data analysis methods but rather in addition to these methods. Together with quantitative analysis of data, such as offered by statistical analysis, image or signal processing, visualization attempts to explore all information inherent in astrophysical data in the most effective way. Data visualization is one aspect of data analysis. Our taxonomy as developed in Section 2 includes identification and access to existing information, preprocessing and quantitative analysis of data, visual representation and the user interface as major components to the software environment of astrophysical data analysis. In pursuing our goal to provide methods and tools for scientific visualization of multi-spectral astrophysical data, we therefore looked at scientific data analysis as one whole process, adding visualization tools to an already existing environment and integrating the various components that define a scientific data analysis environment. As long as the software development process of each component is separate from all other components, users of data analysis software are constantly interrupted in their scientific work in order to convert from one data format to another, or to move from one storage medium to another, or to switch from one user interface to another. We also took an in-depth look at scientific visualization and its underlying concepts, current visualization systems, their contributions and their shortcomings. The role of data visualization is to stimulate mental processes different from quantitative data analysis, such as the perception of spatial relationships or the discovery of patterns or anomalies while browsing through large data sets. Visualization often leads to an intuitive understanding of the meaning of data values and their relationships by sacrificing accuracy in interpreting the data values. In order to be accurate in the interpretation, data values need to be measured, computed on, and compared to theoretical or empirical models (quantitative analysis). If visualization software hampers quantitative analysis (which happens with some commercial visualization products), its use is greatly diminished for astrophysical data analysis. The software system STAR (Scientific Toolkit for Astrophysical Research) was developed as a prototype during the course of the project to better understand the pragmatic concerns raised in the project. STAR led to a better understanding on the importance of collaboration between astrophysicists and computer scientists. Twenty-one examples of the use of visualization for astrophysical data are included with this report. Sixteen publications related to efforts performed during or initiated through work on this project are listed at the end of this report.

Brugel, E. W.↗

FLEET: Flexible Efficient Ensemble Training for Heterogeneous Deep Neural Networks

Parallel training of an ensemble of Deep Neural Networks (DNN) on a cluster of nodes is an effective approach to shorten the process of neural network architecture search and hyper-parameter tuning for a given learning task. Prior efforts have shown that data sharing, where the common preprocessing operation is shared across the DNN training pipelines, saves computational resources and improves pipeline efficiency. Data sharing strategy, however, performs poorly for a heterogeneous set of DNNs where each DNN has varying computational needs and thus different training rate and convergence speed. This paper proposes FLEET, a flexible ensemble DNN training framework for efficiently training a heterogeneous set of DNNs. We build FLEET via several technical innovations. We theoretically prove that an optimal resource allocation is NP-hard and propose a greedy algorithm to efficiently allocate resources for training each DNN with data sharing. We integrate data-parallel DNN training into ensemble training to mitigate the differences in training rates and introduce checkpointing into this context to address the issue of different convergence speeds. Experiments show that FLEET significantly improves the training efficiency of DNN ensembles without compromising the quality of the result.

Guan, Hui↗