Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data processing automation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Methodology for physics-informed generation of synthetic neutron time-of-flight measurement data

Accurate neutron cross section data are a vital input to the simulation of nuclear systems for a wide range of applications from energy production to national security. The evaluation of experimental data is a key step in producing accurate cross sections. There is a widely recognized lack of reproducibility in the evaluation process due to its artisanal nature and therefore there is a call for improvement within the nuclear data community. This can be realized by automating/standardizing viable parts of the process, namely, parameter estimation by fitting theoretical models to experimental data. This automation effort could greatly benefit from a synthetic data resource. This work leverages problem-specific physics, Monte Carlo sampling, and a general methodology for data synthesis to generate unlimited, labelled experimental cross-section data that is statistically indistinguishable to the observed data. Heuristic and, where applicable, rigorous statistical comparisons to observed data support this claim. The demonstration is based on/limited to transmission measurements at Rensselaer Polytechnic Institute (RPI) and energy-differential cross sections in the resolved resonance region (RRR). An open-source software is published alongside this article that executes the complete methodology to produce high-utility synthetic datasets. The goal of this work is to provide an approach and corresponding tool that will allow the evaluation community to begin exploring more data-driven, ML-based solutions to long-standing challenges in the field.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Refining PeakDecoder Version 2

Novel computational tools for processing multidimensional mass spectrometry (MS) data are necessary to enable deeper and automated detection and quantification of metabolites in complex backgrounds. Multidimensional MS data includes measurements from liquid chromatography (LC) and ion mobility spectrometry (IM) separations, and precursor and fragment ion spectra collected in data-independent acquisition (DIA) mode. PeakDecoder is an artificial intelligence (AI)-based software that enables automated interpretation of this kind of data to identify and quantify individual metabolites in complex mixtures. The goal of this project was to improve and re-implement PeakDecoder in a better suited programming language to enable its commercialization.

97 MATHEMATICS AND COMPUTING↗

AI for Nuclear Safeguards Verification

The International Atomic Energy Agency (IAEA) utilizes AI/ML to analyze open-source information, including satellite imagery and scientific publications, to verify the completeness of State declarations regarding nuclear activities. AI/ML already assist the IAEA with automating processes and analysis of large datasets, including satellite imagery and unstructured data, improving efficiency and effectiveness of safeguards implementation. AI/ML in nuclear safeguards come with its own challenges that include the need for large, unbiased datasets, the risk of AI-generated fake information, including the potential for manipulation of satellite imagery.

97 MATHEMATICS AND COMPUTING↗

Smart Semi-Supervised Accumulation of Large Repositories for Industrial Control Systems Device Information

Industrial Control Systems device manufacturers frequently add new features to improve their product performance. Oftentimes, these changes are mainly vendor-driven initiatives, and customers may not be aware of the full impact of these new capabilities on their cybersecurity posture. In the energy sector, this can lead to considerable dissonance between vendor-provided cybersecurity claims and a customer’s responsibility for Operation Technology cybersecurity compliance. Thus, the resulting dynamic verification burden is shifted towards the customer and may pose a significant cybersecurity risk to the energy sector landscape. We found that there is very limited research into cybersecurity auditing for Operational Technology. However, a solution is needed for vetting the vendor-supplied feature claims and their adherence to cybersecurity requirements and standards. We are presently engaged in an effort to develop such a system. This paper demonstrates one vital aspect of this effort in proposing an end-to-end framework to accumulate a large repository of ICS device information for this vetting system, curate the dataset, and conduct extensive processing. This framework is designed to use web scraping, data analytics and Natural Language Processing (NLP) techniques to identify vendor websites, automate the collection of website-accessible documents and automatically derive metadata from them for identification of product documents relevant to the repository. We have found that this automated approach to vendor identification, document extraction into a product repository, and NLP pre-processing is unique and has not been previously presented in the literature. The preliminary work shows that this is feasible and can produce reliable results with minimum supervision. Future work will be built upon this foundation in order to achieve semi-supervised vetting of device technical information – a vital capability for ensuring that vendor-claimed device cybersecurity capabilities match industry requirements.

Ameri, Kimia↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

Probabilistic Modeling of Commercial Building Occupancy Patterns Using Location-Based Map Data: Preprint

Considering occupancy patterns is crucial to simulate buildings' energy use. Current energy models use inputs that simplify the actual diversity in occupancy into static occupancy patterns and are not able to represent the numerous variations in occupancy patterns between buildings and across different locations. Recently, inferring occupancy schedules from metered electricity consumption data was used to model occupancy in commercial buildings. However, the translation from metered data to occupancy schedules requires many assumptions that might not capture the reality, and the process is hindered by the availability of data from advanced metering infrastructure. With the development of information technologies, occupancy modeling should not be limited to traditional approaches. The prevalence of social networks and location services with real-time user feedback provides publicly accessible data via Maps Application Programming Interfaces (APIs) such as Google Maps, SafeGraph, Mapbox, Foursquare, etc. This paper presents an automated framework for modeling parametric occupancy patterns using such APIs to calibrate commercial district buildings' energy models. This process includes three main steps: data extraction and processing, parametric schedules generation, and schedules integration. We demonstrated this framework in districts where we used maps API to generate more accurate behavioral patterns for operations and electric vehicle charging events. We used these patterns to determine differences in energy use across key sociodemographic and spatial parameters. The presented method has the potential for worldwide applications. Users can utilize this framework to extract data for selected locations of interest to create more realistic behavioral patterns for commercial facilities across different districts.

building energy modeling↗

The CanBikeCO Mini Pilot: Preliminary Results and Lessons Learned

In fall 2020, the Colorado Energy Office, as part of the State of Colorado’s “Can Do Colorado” initiative, initiated a project aimed at encouraging energy-efficient transportation during the COVID-19 pandemic. The initial mini-pilot provided e-bikes to 13 low-income households under an individual ownership model. This report assesses the impact of providing this additional mobility option on the travel behavior of participants. It also outlines the lessons learned from deploying a continuous monitoring platform to track the travel behavior. These lessons will influence the evaluation component for the full pilot, which will cover multiple geographic regions, starting in summer 2021, and run for 2 years. The continuous data collection was enabled by a customized version of the open-source e-mission platform, called CanBikeCO, configured with a behavioral gamification feature. The Colorado Energy Office used this system to collect a unique data set consisting of 3 months of partially automated travel diaries, combining sensed and surveyed data and linked with demographic information, from 12 participants. The data collection process worked well overall: users generally liked the app, appreciated the game, and did not complain about battery life. The long tracking period introduced behavioral challenges in user engagement, which we plan to address using repeated patterns and automated status checks for the full pilot. The analysis results, based on the subset of trips with user-reported labels (68%), indicate that the e-bike was the dominant commute mode share (31%), in sharp contrast to the census bicycle commute mode share (<1%). E-bike trips primarily replaced single-occupancy vehicle (SOV) trips (28%), followed closely by walking (24%) and regular bike (20%). The nonmotorized mode replacement corresponds to lower travel time and increased productivity enabled by the program. The emissions impact analysis of the program, computed using trip-level energy intensity factors, indicates savings of 1,367 lbs. of CO 2 . Although the results are strongly positive, the narrow demographic profile of study participants, their limited mobility alternatives, and nonuniform labeling indicate caution in broader interpretation. These preliminary results do suggest that such programs, supported by real-time education and support from program managers, can simultaneously meet equity and sustainability goals. The planned full pilot, addressing the data collection challenges and broadening the geographic scope, will provide additional insights into the generality of this approach.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The CanBikeCO Mini Pilot: Procedure and Preliminary Results

In fall 2020, the Colorado Energy Office, as part of the State of Colorado's "Can Do Colorado" initiative, initiated a project aimed at encouraging energy-efficient transportation during the COVID-19 pandemic. The initial mini-pilot provided e-bikes to 13 low-income households under an individual ownership model. This report assesses the impact of providing this additional mobility option on the travel behavior of participants. It also outlines the lessons learned from deploying a continuous monitoring platform to track the travel behavior. These lessons will influence the evaluation component for the full pilot, which will cover multiple geographic regions, start in summer 2021, and run for 2 years. The continuous data collection was enabled by a customized version of the open-source e-mission platform, called CanBikeCO, configured with a behavioral gamification feature. The Colorado Energy Office used this system to collect a unique data set consisting of 3 months of partially automated travel diaries, combining sensed and surveyed data and linked with demographic information, from 12 participants. The data collection process worked well overall: users generally liked the app, appreciated the game, and did not complain about battery life. The long tracking period introduced behavioral challenges in user engagement, which we plan to address using repeated patterns and automated status checks for the full pilot. The analysis results, based on the subset of trips with user-reported labels (68%), indicate that the e-bike was the dominant commute mode share (31%), in sharp contrast to the census bicycle commute mode share (<1%). E-bike trips primarily replaced single-occupancy vehicle (SOV) trips (28%), followed closely by walking (24%) and regular bike (20%). The non-motorized mode replacement corresponds to lower travel time and increased productivity enabled by the program. The emissions impact analysis of the program, computed using trip-level energy intensity factors, indicates savings of 1,367 lbs. of CO2. Although the results are strongly positive, the narrow demographic profile of study participants, their limited mobility alternatives, and nonuniform labeling indicate caution in broader interpretation. These preliminary results do suggest that such programs, supported by real-time education and support from program managers, can simultaneously meet equity and sustainability goals. The planned full pilot, addressing the data collection challenges and broadening the geographic scope, will provide additional insights into the generality of this approach.

ADVANCED PROPULSION SYSTEMS↗

A systematic feature extraction and selection framework for data-driven whole-building automated fault detection and diagnostics in commercial buildings

In data-driven automated fault detection and diagnostics (AFDD) modeling for building energy systems, feature engineering is a critical process of extracting information from high-dimensional and noisy sensor measurement and turning it into informative and representative inputs or features for data-driven modeling. However, few studies specifically discuss the feature engineering, especially the interactions between feature extraction and feature selection in whole-building AFDD. We developed a systematic feature extraction and selection framework for whole-building AFDD. In this framework, features are aggressively extracted from raw sensor data using statistical feature extraction techniques with various window sizes and statistics. With many features extracted, a hybrid feature selection algorithm that combines the filter and wrapper method then selects the best feature set. The framework considers diversity in the duration of fault behavior among fault types in whole-building AFDD, thus achieving high model generalization. We implemented our developed framework in a virtual testbed calibrated with measured data from Oak Ridge National Laboratory's Flexible Research Platform designed to mimic the operation of a typical small commercial building. The AFDD model is trained by the simulation data generated from the virtual testbed. The results show that (1) the developed framework improves the generalization of the AFDD model by 10.7% compared with literature-reported feature extraction and selection methods and (2) features with diverse window sizes and statistics are selected, providing insight into physical systems beyond the current understanding of buildings and faults and improving the detection and diagnostics of multiple fault types.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Emerging materials intelligence ecosystems propelled by machine learning

We report that the age of cognitive computing and artificial intelligence (AI) is just dawning. Inspired by its successes and promises, several AI ecosystems are blossoming, many of them within the domain of materials science and engineering. These materials intelligence ecosystems are being shaped by several independent developments. Machine learning (ML) algorithms and extant materials data are utilized to create surrogate models of materials properties and performance predictions. Materials data repositories, which fuel such surrogate model development, are mushrooming. Automated data and knowledge capture from the literature (to populate data repositories) using natural language processing approaches is being explored. The design of materials that meet target property requirements and of synthesis steps to create target materials appear to be within reach, either by closed-loop active-learning strategies or by inverting the prediction pipeline using advanced generative algorithms. AI and ML concepts are also transforming the computational and physical laboratory infrastructural landscapes used to create materials data in the first place. Surrogate models that can outstrip physics-based simulations (on which they are trained) by several orders of magnitude in speed while preserving accuracy are being actively developed. Automation, autonomy and guided high-throughput techniques are imparting enormous efficiencies and eliminating redundancies in materials synthesis and characterization. The integration of the various parts of the burgeoning ML landscape may lead to materials-savvy digital assistants and to a human-machine partnership that could enable dramatic efficiencies, accelerated discoveries and increased productivity. Here, we review these emergent materials intelligence ecosystems and discuss the imminent challenges and opportunities. The materials research landscape is being transformed by the infusion of approaches based on machine learning. This Review discusses the emerging materials intelligence ecosystems and the potential of human-machine partnerships for fast and efficient virtual materials screening, development and discovery.

36 MATERIALS SCIENCE↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

An automated integrated web-based smart tool for open stope design

The Stability Graph is a widely used tool for the design of open stopes in underground mining. Many users of the Stability Graph still apply this design method manually. Although the manual approach has benefits, using multiple graphs and stability number computation charts for each stope surface is time-consuming, even for the experienced mining engineer. Current practice in the use of the method also limits data sharing. This paper presents a StopeSoft web-based tool for open stope stability prediction that is developed on the basis of the Stability Graph method and is available at openstope.com. StopeSoft incorporates flexibility in terms of Stability Graph options and incorporates additional critical factors often overlooked. As a web-based tool, StopeSoft encourages and makes data sharing possible globally, focused on expanding the database and improving the current limitations of the Stability Graph to provide practical, reliable solutions for mining engineers, consultants, and academics. The StopeSoft automated process facilitates the process of open stope stability prediction, saving time and minimizing potential human errors. Statistical treatment of the data accounts for the variability of input parameters to emphasize the probabilistic nature of the Stability Graph method. The probabilistic interpretation of the stability states of stope surfaces eliminates the false feeling of absolute stope performance based on its location on the Stability Graph , as implied by the deterministic approach.

58 GEOSCIENCES↗

Beyond Visual Analytics: Human-Machine Teaming for AI-Driven Data Sensemaking

"Detect the expected, discover the unexpected" was the founding principle of the field of visual analytics. This mantra implies that human stakeholders, like a domain expert or data analyst, could leverage visual analytics techniques to seek answers to known unknowns and discover unknown unknowns in the course of the data sensemaking process. We argue that in the era of AI-driven automation, we need to recalibrate the roles of humans and machines (e.g., a machine learning model) as teammates. We posit that by realizing human-machine teams as a stakeholder unit, we can better achieve the best of both worlds: automation transparency and human reasoning efficacy. However, this also increases the burden on analysts and domain experts towards performing more cognitively demanding tasks than what they are used to. In this paper, we reflect on the complementary roles in a human-machine team through the lens of cognitive psychology and map them to existing and emerging research in the visual analytics community. We discuss open questions and challenges around the nature of human agency and analyze the shared responsibilities in human-machine teams.

Sensemaking, human-machine teaming, agency, artifi↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Panorama 360 (Final Report)

This final technical report from the lead institution, USC grant #DE-SC0012636, serves as the final technical report for collaborative institution UNC-CH grant #DE-SC0012390. The goal was to develop a repository and associated capabilities for data collection, ingestion, and analysis for a broad class of DOE applications that span experimental and simulation science workflows. In particular, this work focuses on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: (1) A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); (2) A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; (3) A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and (4) Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

AUTONOMIE VID

Autonomie Vehicle Information Database (VID) offers a comprehensive list of vehicle specifications since 1990. The database details more than 65,000 vehicles with hundreds of attributes. The database is the result of the development of a general automated data collection framework as well as the development of building blocks for processing, cleaning, integrating and analyzing complex data. The data has undergone several layers of outlier detections processes, machine learning based imputations methods have been used to deal with missing data problems, and new fields have been created according to the rules of feature engineering. Thanks to this streamlined data pipelines, the resulting processed aggregated data should deliver a unique level of information to the user in which the content can be efficiently maintained and updated.

Moswd, Ayman↗

DEEP CELLULAR RECURRENT NEURAL ARCHITECTURE FOR EFFICIENT MULTIDIMENSIONAL TIME-SERIES DATA PROCESSING

Efficient processing of time series data is a fundamental yet challenging problem in pattern recognition. Though recent developments in machine learning and deep learning have enabled remarkable improvements in processing large scale datasets in many application domains, most are designed and regulated to handle inputs that are static in time. Many real-world data, such as in biomedical, surveillance and security, financial, manufacturing and engineering applications, are rarely static in time, and demand models able to recognize patterns in both space and time. Current machine learning (ML) and deep learning (DL) models adapted for time series processing tend to grow in complexity and size to accommodate the additional dimensionality of time. Specifically, the biologically inspired learning based models known as artificial neural networks that have shown extraordinary success in pattern recognition, tend to grow prohibitively large and cumbersome in the presence of large scale multi-dimensional time series biomedical data such as EEG. Consequently, this work aims to develop representative ML and DL models for robust and efficient large scale time series processing. First, we design a novel ML pipeline with efficient feature engineering to process a large scale multi-channel scalp EEG dataset for automated detection of epileptic seizures. With the use of a sophisticated yet computationally efficient time-frequency analysis technique known as harmonic wavelet packet transform and an efficient self-similarity computation based on fractal dimension, we achieve state-of-the-art performance for automated seizure detection in EEG data. Subsequently, we investigate the development of a novel efficient deep recurrent learning model for large scale time series processing. For this, we first study the functionality and training of a biologically inspired neural network architecture known as cellular simultaneous recurrent neural network (CSRN). We obtain a generalization of this network for multiple topological image processing tasks and investigate the learning efficacy of the complex cellular architecture using several state-of-the?art training methods. Finally, we develop a novel deep cellular recurrent neural network (CDRNN) architecture based on the biologically inspired distributed processing used in CSRN for processing time series data. The proposed DCRNN leverages the cellular recurrent architecture to promote extensive weight sharing and efficient, individualized, synchronous processing of multi-source time series data. Experiments on a large scale multi-channel scalp EEG, and a machine fault detection dataset show that the proposed DCRNN offers state-of-the-art recognition performance while using substantially fewer trainable recurrent units.

Vidyaratne, Lasitha S.↗

A Modularized Urban Scale Building Energy Modeling Framework Designed with An Open Mind

In recent years, physics-based building energy modeling (BEM) has started being used to evaluate the performance of buildings in the context of connected communities and on an urban scale to study their aggregated energy use, interactions, and impacts on the energy supply infrastructure and environment. The development of urban-scale BEM solutions needs extensive effort. Existing attempts tend to focus on different aspects of BEM on an urban scale, such as collecting as-built building data from different information sources, integrating geometry modeling with geographic information systems (GISs), representing operational and occupancy profiles, automating workflow, processing and visualizing the results, and conducting large-scale simulations. Urban-scale BEM development would benefit from multi-disciplinary research areas and from an open platform to adopt advancements on data sources and tools. For these purposes, this research proposes a modularized bottom-up model creation and simulation framework that is built on the state-of-the-art BEM tools and can accommodate different building stock data. This framework uses a standardized schema to describe building design and operational characteristics, and it can be instantiated from different building survey datasets with heterogeneous structures. The paper demonstrates how thousands of surveyed buildings from the 2012 U.S. Energy Information Administration’s Commercial Buildings Energy Consumption Survey (CBECS) were one-to-one converted to EnergyPlus models through the schema and the model generation process, then simulated with distributed computing, and their results are summarized.

Lei, Xuechen↗