Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data-Driven Day-Ahead PV Estimation Using Autoencoder-LSTM and Persistence Model

Inherent variability in photovoltaic (PV) and associated impacts on power systems is a challenging problem for both the PV owners and the grid operators. Existing statistical and machine learning algorithms typically work well for weather conditions similar to historical data. Furthermore, uncertain weather conditions pose a great challenge to the estimation accuracy of the estimation models. With the enhanced integration of intelligent electronic devices and the realization of associated automation in the power grid, renewable energy data is becoming more accessible, which can be utilized by deep learning models and improve the PV power generation estimation accuracy. In this paper, a hybrid deep learning model driven by external weather data is proposed to do day-ahead PV output forecasting at 15-minute-interval. The proposed model is motivated by the recent advancement of Long-Short-Term-Memory (LSTM) networks and AutoEncoder (AE), which estimates uncertainties in sequence while making the prediction for complex weather conditions. Meanwhile, the persistence model (PM) is used to predict continuous sunny weather conditions. The forecasting result is validated with data from multiple locations

42 ENGINEERING↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Phase evolution and amorphous stability upon solid-state reaction in superlattice like Ge–Sb–Te combinatorial thin-film

In this paper, the superlattice-like (SLL) Ge–Sb–Te combinatorial thin films were prepared by using a high-throughput ion beam sputtering system. The phase evolution and amorphous stability of such films undergoing heat treatment as a function of the coating sequence and modulation period were systematically studied. The composition structure diagram was constructed via an automated process of data obtained by high-throughput synchrotron micro-X-ray diffraction and lab-based micro-X-ray fluorescence. Furthermore, the element distribution and microstructure in the depth direction of the SLL thin films were characterized with time-of-flight secondary ion mass spectrometry and transmission electron microscopy, respectively. These studies demonstrated that the coating sequence has a significant effect on the element distribution in the as-deposited SLL thin films and the structure of the final product upon solid-state reaction. Reducing the modulation period of the SLL thin film improves the stability of the amorphous Ge–Sb–Te phase. This work lays a solid foundation for the rational design of SLL Ge–Sb–Te thin films to improve their performance.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring Secondary Markets to Improve Circularity: A Comparative Case Study of Photovoltaics and Hard-Disk Drives

Each year renewable energy generation increases notably with solar panel installations, but these panels have a limited lifespan and will produce between 2 and 4 million metric tons of waste by 2040. Similarly, there are currently between 20 to 70 million hard-disk drives (HDDs) reaching end-of-life (EOL) annually. The circular economy (CE) strives to recycle and reuse materials that are rare and expensive to obtain, minimizing waste. However, studying the potential circularity of photovoltaics (PV) and HDDs requires various data, for instance, on the maturity of the secondhand markets. In this context, the objective of the present study is to identify the current state of secondhand PV and HDD markets. After conducting a literature review, an automated data collection process was set up for that purpose. The analysis of the literature and collected data assess the maturity of the secondhand PV modules and HDDs markets and highlight differences between them.

agent-based modelling↗

Virtual Inspection of Advanced Manufacturing via Process-Scale Digital Twins (Abbreviated Report)

Inspection and certification comprise the most significant bottlenecks in advanced manufacturing for NNSA applications, often requiring far more time and resources than the fabrication of the parts themselves. Traditional methods, such as manual review and X-ray computed tomography, are not only slow and costly, but also struggle to provide a clear connection between manufacturing instructions and the final performance of critical components. This gap limits both the agility and assurance needed to support the modernization and safety of the United States nuclear stockpile. In response, our Strategic Initiative established a digital twin framework that integrates realtime process monitoring, automated data analysis, and immersive virtual reality collaboration into a unified inspection pipeline. By leveraging data from sensors, machine instructions, and imaging, we created high-fidelity virtual models of manufactured parts that could be rapidly analyzed and certified. This approach was first demonstrated with Direct Ink Write, and then extended to other manufacturing settings, including conventional (or “subtractive”) manufacturing and to predict the end of life performance of parts per the aging and lifetimes programs. The result is a transformational capability: inspection times have been reduced by a factor of 120,000 without loss of accuracy and while simultaneously improving traceability and confidence in part quality. This framework not only streamlines certification for critical applications, but also positions the national security enterprise to respond more flexibly to emerging challenges, supporting agile manufacturing and digital engineering practices across a broad range of mission-relevant domains.

42 ENGINEERING↗

PVAnalytics: A Python Package for Automated Processing of Solar Time Series Data

Multiple publicly available software packages exist that analyze solar time series data, including RdTools and Solar Data Tools, among others. Several of these packages contain their own unique quality assurance (QA) and feature recognition algorithms. The python PVAnalytics package was developed to offer an internally consistent source for these analysis tools, making it easier for the end user to deploy these routines on his or her solar data. The PVAnalytics package currently contains routines for outlier detection, inverter clipping detection, irradiance and temperature checks, orientation checks, and data shift detection, among other functions. These functions have been aggregated from various sources including Solar Forecast Arbiter, RdTools, and the QA process developed by NREL's PV Fleets Initiative. We are continuously adding new functionality to the package, including documentation, examples and algorithms. By bundling QA functionality into a single software package, we hope to make PVAnalytics a comprehensive software library to support analysis of solar metadata and time series data.

data cleaning↗

HyperKube: A Kubernetes Based System for the Automation of Processing and Analysis of Hyperspectral Data Obtained from Multiple Hyperspectral Imaging Systems

Hyperspectral imagery is an emerging field of technology that has enormous potential for remote and proximal sensing in numerous areas of research. The plant phenotyping community is applying this technology to advance the throughput and accuracy of plant phenotypes based on airborne and lab-based hyperspectral imaging technology. Here we report an automated processing and analysis pipeline for four different hyperspectral imaging platforms, discuss the data issues involved, and present a strategy for computing and data architecture to handle hyperspectral data.

47 OTHER INSTRUMENTATION↗

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Automated pipeline processing X-ray diffraction data from dynamic compression experiments on the Extreme Conditions Beamline of PETRA III

Presented and discussed here is the implementation of a software solution that provides prompt X-ray diffraction data analysis during fast dynamic compression experiments conducted within the dynamic diamond anvil cell technique. It includes efficient data collection, streaming of data and metadata to a high-performance cluster (HPC), fast azimuthal data integration on the cluster, and tools for controlling the data processing steps and visualizing the data using the DIOPTAS software package. This data processing pipeline is invaluable for a great number of studies. The potential of the pipeline is illustrated with two examples of data collected on ammonia–water mixtures and multiphase mineral assemblies under high pressure. The pipeline is designed to be generic in nature and could be readily adapted to provide rapid feedback for many other X-ray diffraction techniques, e.g. large-volume press studies, in situ stress/strain studies, phase transformation studies, chemical reactions studied with high-resolution diffraction etc.

97 MATHEMATICS AND COMPUTING↗

Informing forest carbon inventories under the Paris Agreement using ground-based forest monitoring data

Human interactions with forests have shaped Earth's climate for millennia and will continue to do so as we target net-zero emission goals. Accurately characterizing these climate impacts requires making reliable forest carbon data available for forest monitoring and planning. Here, we develop a semi-automated process for submitting forest carbon measurements from the largest relevant scientific database to the International Panel on Climate Change's Emission Factor Database, which currently has sparse forest carbon data. Building this bridge from scientific research to international policy is an important step towards managing forests in a net-zero motivated future. Humans have been influencing Earth's climate via transformative impacts on forests for millennia, and forests are now recognized as critical to climate change mitigation under the Paris Agreement. The efficacy of climate change mitigation planning and reporting depends on quality data on forest carbon (C) stocks and changes. The Emission Factor Database (EFDB) of the International Panel on Climate Change (IPCC) is intended to be a definitive source for such data, but needs comprehensive and well-documented data to be so. To facilitate submission of forest C estimates from scientific studies to EFDB, we develop and document a process for semi-automated data submission from the Global Forest C database (ForC v4.0), which is the largest compilation of ground-based forest C estimates. We then assess the data currently available through ForC and provide recommendations for improving forest data collection, analysis, and reporting. As of September 2024, ForC contained ~19,286 records potentially relevant to EFDB, 1068 of which had been submitted and posted to EFDB. These represented 19% of the total EFDB records for forest land. Records were unevenly distributed across variables and geographic regions. ForC records (37%) reviewed could not be submitted because the original publication lacked required information. In the future, ground-based forest C estimates should target gaps in the record, and studies should ensure that they report all information necessary for inclusion in EFDB. Given that climate change is rapidly impacting the world's forests, timely reporting of recent estimates will be critical to accurate forest C inventories.

54 ENVIRONMENTAL SCIENCES↗

In Situ Transmission Electron Microscopy: Signal processing challenges and examples

Transmission electron microscopy (TEM) is a powerful tool for imaging material structure and characterizing material chemistry. Recent advances in data collection technology for TEM have enabled high-volume and high-resolution data collection at a microsecond frame rate. Here, taking advantage of these advances in data collection rates requires the development and application of data processing tools, including image analysis, feature extraction, and streaming data processing techniques. In this article, we highlight a few areas in materials science that have benefited from combining signal processing and statistical analysis with data collection capabilities in TEM and present a future outlook on opportunities of integrating signal processing with automated TEM data analysis.

36 MATERIALS SCIENCE↗

BAMCensus (The Behavior and Advanced Mobility Census Dataset Aggregator) [SWR-25-120]

This software is a high-performance tool developed in Rust for downloading and processing large-scale geospatial datasets, specifically focusing on US Census data. It is designed to address scaling limitations found in existing tools, such as R's [tidycensus](https://walker-data.com/tidycensus/), by providing performant streaming dataset JOIN operations between various US Census datasets (like ACS and LEHD) and their corresponding geometries stored on the TIGER/Lines web server. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. Its primary motivation stems from the need for a high-performance solution to combine spatial datasets with graph traversals within the context of mobility analysis tooling being developed at NREL's Behavior and Advanced Mobility (BAM) group.

Fitzgerald, Robert [National Renewable Energy Labo↗

Oscilloscope Data Push Program

Data acquisition (DAQ) is a complex and costly process. Creating DAQ systems for analyzing a system requires expensive electronics and a dedicated team of engineers for support, posing a challenge for users who readily need data. This project is a proof of concept to create a temporary or one-off DAQ system using equipment commonly available to every team. We aim to automate the data acquisition process from the Rohde \& Schwarz RTO 1044 oscilloscope, convert the acquired binary data into floating point values, and store the results in a CSV file format. By developing a Python program to handle these tasks, we seek to reduce the manual effort involved in data collection, significantly increasing efficiency.

Osei-Tutu, Jason↗

Challenges and Opportunities for Electric Utility Modeling and Asset Valuation Frameworks: Case Study on Valuing New Pumped Storage Hydropower

Asset valuation by electric utilities is becoming increasingly difficult in the rapidly changing electric sector. Rapid deployment of variable generation and inverter-based storage systems along with uncertain demand growth, climate, policies, and other factors create a challenging environment for understanding the value proposition of a new potential asset. This report describes an effort between the Tennessee Valley Authority (TVA) and three U.S. Department of Energy laboratories to perform a detailed review of utility modeling and analysis practices for asset valuation and identify challenges and opportunities for advancing its methods into the future. It focuses on a case study of new potential pumped storage hydropower (PSH) because of growing interest in new PSH capacity to provide energy balancing, firm capacity, and a range of ancillary services. Staff from the DOE labs conducted systematic interviews about current practices in capacity expansion modeling, production-cost modeling, hydrological modeling, and transmission stability modeling while also discussing how scenario analysis is conducted and how models and data are integrated. The effort resulted in a set of model, integration, and scenario recommendations that could be valuable to TVA, other utilities, system operators, and other stakeholders conducting integrated grid analysis. Individual model recommendations suggest exploring computational tradeoffs with detail and resolution across spatiotemporal structure, supply- and demand-side details, transmission overlays, market interactions, and ancillary services. Automated processes to pass data between models and conduct larger scenario suites could also enhance valuation practices by enabling a more consistent study of asset value across a broader range of uncertain future grid conditions where PSH could be particularly valuable. TVA and other industry stakeholders can learn from and adapt applied research-grade methods developed by DOE laboratories and other research institutions to improve decision making and accelerate progress towards a reliable, economic, sustainable energy system.

13 HYDRO ENERGY↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

libwfa: Wavefunction analysis tools for excited and open‐shell electronic states

Abstract An open‐source software library for wavefunction analysis, libwfa, provides a comprehensive and flexible toolbox for post‐processing excited‐state calculations, featuring a hierarchy of interconnected visual and quantitative analysis methods. These tools afford compact graphical representations of various excited‐state processes, provide detailed insight into electronic structure, and are suitable for automated processing of large data sets. The analysis is based on reduced quantities, such as state and transition density matrices (DMs), and allows one to distill simple molecular orbital pictures of physical phenomena from intricate correlated wavefunctions. The implemented descriptors provide a rigorous link between many‐body wavefunctions and intuitive physical and chemical models, for example, exciton binding, double excitations, orbital relaxation, and polyradical character. A broad range of quantum‐chemical methods is interfaced with libwfa via a uniform interface layer in the form of DMs. This contribution reviews the structure of libwfa and highlights its capabilities by several representative use cases. This article is categorized under: Software > Quantum Chemistry Theoretical and Physical Chemistry > Spectroscopy

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

SolarAPP+ Performance Review (2022 Data)

The Solar Automated Permit Processing Plus (SolarAPP+) platform is an online portal to facilitate and expedite rooftop solar photovoltaic (PV) permitting processes. SolarAPP+ allows PV contractors to upload system specifications, have those specifications automatically reviewed for code compliance, and receive instant approval for code-compliant systems. SolarAPP+ also provides inspection checklists to verify installation practices and adherence to approved designs. SolarAPP+ is available to authorities having jurisdiction (AHJs) at no cost. This report is part of an ongoing series of reviews of SolarAPP+ performance. Consistent with previous performance reviews, we summarize SolarAPP+ adoption trends to date and compare various metrics for PV systems permitted through SolarAPP+ versus systems permitted through conventional AHJ permitting processes. As of the end of 2022, the National Renewable Energy Laboratory (NREL) had contacted over 1,500 AHJs with significant solar permitting volume regarding SolarAPP+. Of those, 607 AHJs had at least expressed interest in the platform. 16 AHJs had begun piloting the platform and 15 of these had publicly launched the platform by the end of 2022. In 2022, 206 installers submitted 11,092 permits through the SolarAPP+ platform, including 708 permits for solar+storage systems. SolarAPP+ permits accounted for around 37% of all permits issued in participating AHJs. We compare permitting timelines through SolarAPP+ to traditional AHJ permitting processes to assess the platform's performance. Consistent with previous SolarAPP+ performance reviews, we find that permitting timelines are significantly shorter for SolarAPP+ projects. Based on median timelines, a typical SolarAPP+ project is permitted and inspected 8 business days sooner than traditional projects. We estimate that automatic SolarAPP+ permitting saved between 3,500 and 13,900 hours of AHJ staff time in 2022. Finally, we find evidence that SolarAPP+ may improve inspection outcomes, with SolarAPP+ projects failing inspections about 28% less frequently than traditional projects.

14 SOLAR ENERGY↗

SolarAPP+ Performance Review (2023 Data)

The Solar Automated Permit Processing Plus (SolarAPP+) platform is an online portal to facilitate and expedite rooftop solar photovoltaic (PV) and battery storage permitting processes. SolarAPP+ allows PV contractors to upload system specifications, have that information automatically reviewed for code compliance, and receive instant approval for code-compliant systems, reducing authority having jurisdiction (AHJ) staff time needed for review. SolarAPP+ also provides inspection checklists to verify installation practices and adherence to approved designs. SolarAPP+ is available to AHJs at no cost. This report is part of an ongoing series of reviews of SolarAPP+ performance. Consistent with previous performance reviews, we summarize SolarAPP+ adoption trends to date and compare various metrics for PV systems permitted through SolarAPP+ versus systems permitted through traditional AHJ permitting processes. As of the end of 2023, the National Renewable Energy Laboratory (NREL) had contacted over 1,700 AHJs with significant solar permitting volume regarding SolarAPP+. Of those, 793 AHJs had expressed interest in the platform as of the end of 2023. 161 AHJs had begun piloting the platform and 97 of these had publicly launched the platform by the end of 2023. In 2023, 668 installers submitted 18,906 permits through the SolarAPP+ platform, including 4,834 permits submitted as part of a solar plus storage program. SolarAPP+ permits accounted for around 43% of all permits issued in participating AHJs. We compare permitting timelines through SolarAPP+ to traditional AHJ permitting processes to assess the platform's performance. Consistent with previous SolarAPP+ performance reviews, we find that permitting timelines are significantly shorter for SolarAPP+ projects. Based on median timelines, a typical SolarAPP+ project is permitted and inspected 14.5 business days sooner than traditional projects. We estimate that automatic SolarAPP+ permitting saved around 7,200 hours of AHJ staff time in 2023. Finally, we estimate that SolarAPP+ eliminated over 150,000 business days in permitting-related delays in 2023.

14 SOLAR ENERGY↗