Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “preprocessing software”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

25 records · Page 2

Tutorial: Machine-Learning-Based CREASE-2D Analysis of 2D SAXS Profiles to Characterize Anisotropic Nanostructures in Soft Materials

We present a tutorial to guide users on how to extend the Computational Reverse Engineering Analysis of Scattering Experiments-2D (CREASE-2D) framework to interpret their experimental two-dimensional small-angle scattering (SAS) data from soft materials (e.g., polymers, peptide amphiphiles, biomolecular fibrils). Unlike most traditional SAS analysis approaches, which typically rely on azimuthally averaged onedimensional (1D) profiles, CREASE-2D utilizes the complete 2D scattering profile to reveal information about anisotropy in the structure. In past applications, CREASE has provided insights into complex structural features, including the cross-sectional shapes of assembled nanostructures and dispersity in these features, which are difficult to discern with existing analytical models. While (1D- ) CREASE has been applied to SANS and SAXS data, this tutorial shares the steps for implementing CREASE-2D using an example of a dipeptide solution system, for which we have SAXS data. We present details for these steps involved in using CREASE-2D to interpret SAXS profiles: how to preprocess SAXS data, define relevant structural features, generate three-dimensional real-space structures for specific values of these features, train a machine learning (ML) surrogate model to predict scattering profiles for given structural features, and optimize these features using genetic algorithms (GA). Then, we use these steps to interpret complex 2DSAXS data collected from dipeptide solutions that, in microscopy images, exhibit nanoscale structures that could be elliptical tubes/ flat tapes/cylinders or a combination of these cross sections. Open-source codes, computational hardware, and software requirements, as well as the strengths and limitations of this protocol, are also presented. We expect researchers working with (soft) biomaterials, peptide amphiphiles, amphiphilic polymer solutions, polymer nanocomposites, and blends of particles/polymers will find this CREASE-2D method and this tutorial of use.

CREASE↗

An Envelope Time Synchronous Averaging for Wind Turbine Gearbox Fault Diagnosis

Vibration-based condition monitoring techniques are widely used for diagnosing faults in rotating machines. These techniques are implemented in the time domain, the frequency domain, or both. However, the composite and noisy nature of the raw data collected requires a preprocessing stage such as filtering and decomposition using in-depth processing techniques. Moreover, these methods require good frequency resolution and involve examining a broad frequency range to discern both healthy and faulty cases. In this work, we introduce a simple and fast diagnostic scheme for wind turbine gear teeth wear based on time domain analysis. The proposed method is based on the local minima interpolation of a filtered version of the vibration signal following time synchronous averaging (TSA) technique. Given tachometer signal, the TSA of the vibration data is performed using MTALAB software. Then, local minima of the filtered signal are interpolated using the Piecewise Cubic Hermite Interpolating Polynomial (PCHIP) function. The variance of the interpolated curve built a gear fault index. The derived fault index resulting of the proposed technique allows a substantial distinction between the healthy and faulty cases. Its efficiency is validated using 10 real-world datasets of vibration stemmed from a wind turbine planetary gearbox. The proposed method boasts a low computation time and ease of interpretation, specifically beneficial for gearbox fault diagnosis purposes.

fault diagnosis↗

Phasor-Measurement-Unit-Based Data Analytics Using Digital Twin and PhasorAnalytics Software

A major objective of this project was to apply GE’s commercial machine learning and data analytics toolsets to large-scale, real-world, anonymized Phasor Measurement Unit (PMU) datasets in order to extract signatures, correlated and/or causal factors, and precursor patterns associated with significant power system phenomena. The project had a particular emphasis on extraction of insights relevant to asset health monitoring, real-time load modeling and cybersecurity monitoring. Additionally, the team was directed to undertake a comprehensive data quality analysis for the provided datasets and encouraged to estimate the ‘machine-learning readiness’ of the datasets by documenting any major obstacles to the application of commercial machine learning algorithms. To accomplish the aforementioned objectives, the project team’s work centered around the identification of key event signatures and application of the identified event signatures for event detection and event classification. The industry-validated, semi-supervised machine learning strategy employed for event signature identification involved several major tasks, including data-preprocessing, generation of an overabundance of features, normal data identification, normality modeling, and event signature identification through a methodical, quantitative ranking of features in order of relevance to each studied event type. Throughout the project, data quality issues and mitigation techniques were investigated. In this report, insights are provided regarding the readiness of the provided synchrophasor datasets for application of machine learning and data analytics. The methodologies employed for this technical strategy are summarized in this report. With regards to data preprocessing and feature generation, the provided Training and Test Datasets were ingested into GE’s big data environment. Subsequently, the team applied bad data cleansing and data imputation scripts, event detection scripts, and application programming interfaces (APIs) to the datasets for convenient data access. The project team completed development and validation of dozens of physics-based, statistics-based and transformation-based feature functions used for the extraction of over 60 synchrophasor features. Using a new parallel feature generation technology developed on this project, over 60 features have been rapidly generated for the full two years’ worth of Training and Test Dataset data associated with both the Eastern and Western interconnects. Even accommodating for temporal down-sampling inherent to the feature extraction procedure, this parallel feature generation activity resulted in a massive feature set with a storage requirement approximately equal to that of the raw training dataset itself. With regards to normal data identification and normality modeling, a normality model was built using the feature data extracted from the Training Dataset and iteratively refined subsequent to incremental adjustments and expansions of the Training Dataset feature data. With respect to event characterization and signature identification, an event signature identification pipeline was developed and used in conjunction with the normality model to identify over 15 event signatures for key event categories within the Training Dataset. The identified event signatures were used to characterize hundreds of key events in terms of relative severity, duration, and location of the event. An investigation was undertaken to identify correlated and causal factors involved in transformer events. A separate investigation into temporal trends in ring-down analysis results was undertaken to determine possible associations between system dynamics and various other factors such as loading, season or year. To validate the identified event signatures, additional work was undertaken to develop signature-based anomaly detection and classification tools suitable for convenient application to the synchrophasor datasets. The anomaly detection and classification tools, suitable for online application, were then applied to the entirety of the Eastern Interconnect Training and Test Datasets. Performance of the event detection and classification tools was evaluated upon receipt of the Test Dataset event logs (i.e., the labels for events contained in the Test Dataset), and promising results were obtained despite several challenges (documented herein) associated with application of supervised or semi-supervised machine learning methods to large-scale, anonymized datasets. Finally, the detection and classification tools were used to detect, classify, and characterize thousands of new events not included in the original event logs provided by the DOE within both the Training and Test Datasets.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Volumetric Rendering on Wavelet-Based Adaptive Grid

Numerical modeling of physical phenomena frequently involves processes across a wide range of spatial and temporal scales. In the last two decades, the advancements in wavelet-based numerical methodologies to solve partial differential equations, combined with the unique properties of wavelet analysis to resolve localized structures of the solution on dynamically adaptive computational meshes, make it feasible to perform large-scale numerical simulations of a variety of physical systems on a dynamically adaptive computational mesh that changes both in space and time. Volumetric visualization of the solution is an essential part of scientific computing, yet the existing volumetric visualization techniques do not take full advantage of multi-resolution wavelet analysis and are not fully tailored for visualization of a compressed solution on the wavelet-based adaptive computational mesh. Our objective is to explore the alternatives for the visualization of time-dependent data on space-time varying adaptive mesh using volume rendering while capitalizing on the available sparse data representation. Two alternative formulations are explored. The first one is based on volumetric ray casting of multi-scale datasets in wavelet space. Rather than working with the wavelets at the finest possible resolution, a partial inverse wavelet transform is performed as a preprocessing step to obtain scaling functions on a uniform grid at a user-prescribed resolution. As a result, a solution in physical space is represented by a superposition of scaling functions on a coarse regular grid and wavelets on an adaptive mesh. An efficient and accurate ray casting algorithm is based just on these coarse scaling functions. Additional details are added during the ray tracing by taking an appropriate number of wavelets into account based on support overlap with the interpolation point, wavelet coefficient magnitude, and other characteristics, such as opacity accumulation (front to back ordering) and deviation from frontal viewing direction. The second approach is based on complementing of wavelet-based adaptive mesh to the traditional Adaptive Mesh Refinement (AMR) mesh. Both algorithms are illustrated and compared to the existing volume visualization software for Rayleigh-Benard thermal convection and electron density data sets in terms of rendering time and visual quality for different data compression of both wavelet-based and AMR adaptive meshes.

Vezolainen, Alexei V.↗

Model Data Archive Associated with Manuscript "Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon"

This data package supports the publication “Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon” by Li et al. (2026). The package contains processed model inputs, configuration files, restart files, simulation outputs, scripts, and visualization products used to evaluate post-fire dissolved organic carbon (DOC) dynamics in the Naches River Watershed, Washington, USA, following the 2021 Schneider Springs Fire. The modeling workflow couples ELM-BGC, the biogeochemistry-enabled Energy Exascale Earth System Model Land Model; ATS, the Advanced Terrestrial Simulator for integrated surface-subsurface hydrology; and PFLOTRAN, a reactive transport model for multicomponent aqueous geochemistry. Together, these models simulate how wildfire-induced changes in vegetation, litter, coarse woody debris, and soil organic matter influence DOC production, transport, and reaction from burned hillslopes to stream networks. The archive includes preprocessed meteorological, geospatial, hydrologic, and biogeochemical forcing data; ELM-BGC-derived DOC source terms; ATS mesh files; PFLOTRAN reactive-transport inputs; model configuration files; spin-up and transient restart files; watershed-scale diagnostic outputs; stream concentration time series; and figures or visualization files used to inspect and reproduce key results. File types include Hierarchical Data Format 5 (HDF5) files for gridded forcing and model-coupling data, model input and configuration files for ELM-BGC, ATS, and PFLOTRAN, restart and simulation-output files generated by the modeling workflow, tabular or time-series diagnostic outputs, scripts for post-processing and figure generation, and image or visualization products associated with the manuscript. Use of the package depends on the intended task. Re-running the simulations requires the relevant modeling software, including ELM-BGC, ATS, and PFLOTRAN as ATS's geochemical engine. Inspecting outputs and reproducing figures requires Python with scientific plotting libraries such as Matplotlib, and three-dimensional model outputs may be viewed with ParaView. Geographic information system files or maps may be inspected with ArcGIS Pro or comparable GIS software. The data package is intended to enable traceability, reuse, and partial reproduction of the coupled land-to-watershed hydro-biogeochemical modeling workflow used to test how wildfire disturbance affects terrestrial carbon pools and downstream DOC dynamics.

ATS↗

Classification of River Catchments in the Contiguous United States: Code, Dataset, Similarity Patterns, and Resulting Classes

This dataset serves as supplementary information for the paper by Ciulla F. and Varadharajan C. A Network Approach for Multiscale Catchment Classification using Traits (see reference 1). It contains environmental and physical catchment traits, such as temperatures, precipitation, land use and human interference, from 9067 sites across the contiguous United States (CONUS). The purpose of this dataset is to provide information for a better trait-based categorization of river catchments in the CONUS using networks as an analytical tool. The traits variables match the ones present in the GAGES-II dataset and the preprocessing steps are described in the Methods section (processed_dataset.csv). Additionally we include the topologies (nodes, edges and clusters, also referred as classes) of the catchment network and traits network generated by said dataset (csv and json files). A series of tables support the information carried by the network providing more detailed descriptions of cluster components (SI1.pdf). A summary of all the plots of clusters of catchments with at least 50 nodes is provided (SI2.pdf). The characteristic traits for each cluster of catchments is presented as z-score (traits_categories_zscores_per_catchment_class.csv). The link to the hydrological behavior of clusters of catchments is displayed by boxplots, each describing a particular river discharge index (SI3.pdf). Both csv and json files can be read by common text editors but the data contained into them can be better handled using programming languages like python and database oriented libraries like pandas. Pdf files can be read by any pdf reader software.[02-23-2024] Update: The code and datasets necessary to reproduce the results of the study are available as a zipped repository (code_datasets_catchments_similarity.zip).

54 ENVIRONMENTAL SCIENCES↗

Real-Time Optimization Workflow Status Update

Economically optimal and safe operation of integrated energy systems (IES) requires optimization at many different time scales. A real-time optimization (RTO) workflow will attempt to maximize revenue and minimize operational costs on a time scale of minutes to hours. Such a workflow requires the use of a digital twin (DT), which is a virtual representation of a physical system. The DT is updated using real-time data from the physical system, and serves as a model in an optimization framework. The optimization results are then sent back to the physical system to complete the loop. This report details the progress made in developing building blocks for a DT/RTO framework. The Risk Analysis Virtual Environment (RAVEN) platform within the Framework for Optimization of Resources and Economics (FORCE) tool suite can perform many of the tasks required for building a DT and performing RTO. The first item of this report details RAVEN enhancements that enable RAVEN workflows to be run in various environments. Data communication between the physical system and its DT is essential for successful RTO. This includes preprocessing real-time data, loading data into a data warehouse, and querying the stored data. The second section of this report describes the progress made in implementing an adapter in Python in order for Deep Lynx to handle the data communication. Typical dispatch optimization frameworks are built on linear programming (LP). The prototype RTO workflow developed in this report uses an LP problem as a part of a receding-horizon- or economic model predictive control (EMPC) based optimization. The third section of this report details the framework of an RTO workflow in which the system consists of a simple electrical storage device. A DT can be built from a reduced-order model (ROM). Integrating a ROM into a typical LP optimization framework has been challenging because most optimization packages require the user to write algebraic expressions for the system model. The final section of this report shows how an externally built RAVEN ROM can be integrated in an RTO framework by using the Python package Pyomo. This demonstrates the RTO workflow capability from a software-only perspective and is an important step in demonstrating the capability to implement an RTO workflow for a physical system.

97 MATHEMATICS AND COMPUTING↗