Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “online algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

A public website for the automated assessment and validation of SARS-CoV-2 diagnostic PCR assays

Abstract Summary Polymerase chain reaction-based assays are the current gold standard for detecting and diagnosing SARS-CoV-2. However, as SARS-CoV-2 mutates, we need to constantly assess whether existing PCR-based assays will continue to detect all known viral strains. To enable the continuous monitoring of SARS-CoV-2 assays, we have developed a web-based assay validation algorithm that checks existing PCR-based assays against the ever-expanding genome databases for SARS-CoV-2 using both thermodynamic and edit-distance metrics. The assay-screening results are displayed as a heatmap, showing the number of mismatches between each detection and each SARS-CoV-2 genome sequence. Using a mismatch threshold to define detection failure, assay performance is summarized with the true-positive rate (recall) to simplify assay comparisons. Availability and implementation The assay evaluation website and supporting software are Open Source and freely available at https://covid19.edgebioinformatics.org/#/assayValidation, https://github.com/jgans/thermonucleotide BLAST and https://github.com/LANL-Bioinformatics/assay_validation. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Introduction to ALERT: A User-Interface Tool for Resilient Optimization

As the cyber-physical systems grow in complexity, there is a need for proactive resilience strategies that involve online, adaptive control actions to best prepare for any impending adversarial events. In this technical effort, supported by RD2C LDRD initiative, the project team designed and demonstrated online strategies – referred to as ALERT controls – for proactive and adaptive tuning of existing optimal controls in a microgrid, with quantifiably assured margins of resilience to various cyber-physical adversarial events. This ALERT functionality is made available to the end-users, e.g., the system operators, via an interactive user-interface. The end-users will not only be able to use the interface to visualize the system’s operation under various cyber-physical adversarial scenarios, but also evaluate the amount of tolerance the system has against selected adversarial perturbations of interest (e.g., malfunctioning sensors, suspected attacked measurements) via the adversarial plots. In this technical report, we briefly outline the algorithmic modules of the developed ALERT control technology, and introduce the user-interface tool that allows end-users (e.g., microgrid operators) to enter their system description, specify various operational and resilience requirements, and evaluate the impact of the control decisions via illustrative plots.

97 MATHEMATICS AND COMPUTING↗

Machine-Learning Enabled Evaluation of Probability of Piping Degradation In Secondary Systems of Nuclear Power Plants

The transition to condition-based, risk-informed automated maintenance will contribute to a significant reduction of operations and maintenance costs that account for the majority of nuclear power generation costs. Furthermore, of the operations and maintenance costs in U.S. plants, approximately 80% are labor costs. To address the issue of rising operating costs and economic viability, technologies used to perform online monitoring of piping and other secondary system structural components in commercial nuclear power plants (NPPs) are under evaluation. These online monitoring systems have the potential to identify when a more detailed inspection is needed using real time measurements, rather than at a pre-determined inspection interval thus reducing the maintenance cost. This paper describes distributed high-temperature stable fiber sensors fabricated in optical fibers through a roll-to-roll laser direct writing process using femtosecond lasers. Using phase-sensitive optical time domain reflectometry, distributed acoustic and vibration sensors can be developed and deployed to critical components and systems in NPPs to perform active measurements with spatial resolution down to 0.5-meter throughout the piping systems. Complex acoustic and vibration signatures harnessed by distributed sensors are registered and analyzed by artificial intelligence algorithms for degradation detection and flaw identification. Piping elbows with machined-in flaws were instrumented with fiber sensors. High-spatial-resolution data were used to develop and validate machine learning algorithms, including both linear and nonlinear regression, and classification. Additionally, classification and sensor analysis were also performed for data analysis. The paper concludes with recommendations and future work on applications of machine learning enabled high-resolution fiber sensors for piping degradation monitoring in current or future NPPs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Heterogeneous data-processing optimization with CLARA’s adaptive workflow orchestrator

The hardware landscape used in HEP and NP is changing from homogeneous multi-core systems towards heterogeneous systems with many different computing units, each with their own characteristics. To achieve maximum performance with data processing, the main challenge is to place the right computing on the right hardware. In this paper, we discuss CLAS12 charge particle tracking workflow orchestration that allows us to utilize both CPU and GPU to improve the performance. The tracking application algorithm was decomposed into micro-services that are deployed on CPU and GPU processing units, where the best features of both are intelligently combined to achieve maximum performance. In this heterogeneous environment, CLARA aims to match the requirements of each micro-service to the strength of a CPU or a GPU architecture. A predefined execution of a micro-service on a CPU or a GPU may not be the most optimal solution due to the streaming data-quantum size and the data-quantum transfer latency between CPU and GPU. So, the CLARA workflow orchestrator is designed to dynamically assign micro-service execution to a CPU or a GPU, based on the online benchmark results analyzed for a period of real-time data-processing.

Gyurjyan, Vardan↗

Machine Learning (ML) Classifier to Assist Metadata Creation

The Atmospheric Radiation Measurement (ARM) Data Center is responsible for the timely collection, archival, and curation of science data products. These products are freely available through an online data repository. Metadata creation is paramount for scientific users to find and access over seven petabytes of atmospheric science data. The hierarchical metadata structure allows users to search for information at both broad and narrow levels. This project aims to leverage 30 years’ worth of manually created metadata to enable machine predictions of broad-term classifications from narrow-term descriptions. These classification predictions would assist metadata coordinators with their term selections. This paper discusses the cleaning and preprocessing of the training data, the pipeline developed to determine the best model for this task, and the creation of an API metadata classifier for ARM measurement metadata. Our results show that the Linear Support Vector Classification (LinearSVC) algorithm, along with the Term Frequency – Inverse Document Frequency (TF-IDF) vectorizer, is well-suited for our multi-class classification task. Lengthier input training data led to better results, and artificial balancing was unnecessary for this particular use case. This predictive classifier enhances efficiency in metadata creation, as well as supports greater consistency and accuracy in metadata tagging.

Collier, Hannah [ORNL] (ORCID:0000000341284292)↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine

03 NATURAL GAS↗

Machine learning on FPGA for event selection

Real-time data processing is a frontier field in experimental particle physics. The application of FPGAs at the trigger level is used by many current and planned experiments (CMS, LHCb, Belle2, PANDA). Usually they use conventional processing algorithms. LHCb has implemented Machine Learning (ML) elements for real-time data processing with a triggered readout system that runs most of the ML algorithms on a computer farm. The work described in this article aims to test the ML-FPGA algorithms for streaming data acquisition. Herein, there are many experiments working in this area and they have a lot in common, but there are many specific solutions for detector and accelerator parameters that are worth exploring further. This report describes the purpose of the work and progress in evaluating the ML-FPGA application.

47 OTHER INSTRUMENTATION↗

CONSTAX2: improved taxonomic classification of environmental DNA markers

Abstract Summary CONSTAX—the CONSensus TAXonomy classifier—was developed for accurate and reproducible taxonomic annotation of fungal rDNA amplicon sequences and is based upon a consensus approach of RDP, SINTAX and UTAX algorithms. CONSTAX2 extends these features to classify prokaryotes as well as eukaryotes and incorporates BLAST-based classifiers to reduce classification errors. Additionally, CONSTAX2 implements a conda-installable command-line tool with improved classification metrics, faster training, multithreading support, capacity to incorporate external taxonomic databases and new isolate matching and high-level taxonomy tools, replete with documentation and example tutorials. Availability and implementation CONSTAX2 is available at https://github.com/liberjul/CONSTAXv2, and is packaged for Linux and MacOS from Bioconda with use under the MIT License. A tutorial and documentation are available at https://constax.readthedocs.io/en/latest/. Data and scripts associated with the manuscript are available at https://github.com/liberjul/CONSTAXv2_ms_code. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data (Final Report)

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components: (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine.

20 FOSSIL-FUELED POWER PLANTS↗

Deep Learning Image Segmentation for Atmospheric Rivers

Abstract The identification of atmospheric rivers (ARs) is crucial for weather and climate predictions as they are often associated with severe storm systems and extreme precipitation, which can cause large impacts on society. This study presents a deep learning model, termed ARDetect, for image segmentation of ARs using ERA5 data from 1960 to 2020 with labels obtained from the TempestExtremes tracking algorithm. ARDetect is a convolutional neural network (CNN)-based U-Net model, with its structure having been optimized using automatic hyperparameter tuning. Inputs to ARDetect were selected to be the integrated water vapor transport (IVT) and total column water (TCW) fields, as well as the AR mask from TempestExtremes from the previous time step to the one being considered. ARDetect achieved a mean intersection-over-union (mIoU) rate of 89.04% for ARs, indicating its high accuracy in identifying these weather patterns and a superior performance than most deep learning–based models for AR detection. In addition, ARDetect can be executed faster than the TempestExtremes method (seconds vs minutes) for the same period. This provides a significant benefit for online AR detection, especially for high-resolution global models. An ensemble of 10 models, each trained on the same dataset but having different starting weights, was used to further improve on the performance produced by ARDetect, thus demonstrating the importance of model diversity in improving performance. ARDetect provides an effective and fast deep learning–based model for researchers and weather forecasters to better detect and understand ARs, which have significant impacts on weather-related events such as floods and droughts.

Galea, Daniel↗

Monitoring Noble Gases (Xe and Kr) and Aerosols (Cs and Rb) in a Molten Salt Reactor Surrogate Off-Gas Stream Using Laser-Induced Breakdown Spectroscopy (LIBS)

In this study with surrogate materials we show that laser-induced breakdown spectroscopy (LIBS) is a robust tool with promising capability toward monitoring gaseous (Xe and Kr) and aerosol (Cs and Rb) species in an off-gas stream from a molten salt reactor (MSR). MSRs will continually evolve fission products into the cover gas flowing across the reactor headspace. The cover gas entrains Xe and Kr gases, along with aerosol particles, before passing into an off-gas treatment system. Univariate models of Xe and Kr peaks showed a strong correlation to concentration indicated by their coefficients of determination of 0.983 and 0.997, respectively. Multivariate models were built for all four analytes using partial least squares regression coupled with preprocessing steps including normalization, trimming, and/or genetic algorithm derived filters. The models were evaluated by predicting the concentrations of the analytes in four validation samples, in which all calibration models were successfully validated at a confidence interval of 99.9%. Finally, pressure controllers were used to regulate the mass flow rate of Kr flowing into the measurement cell in sinusoidal and stepwise waveforms to test the real-time monitoring capabilities of the regression models. Both univariate and partial least squares Kr models were able to successfully quantify the gas concentration in the real-time evaluation. The root mean squared error of prediction (RMSEP) values for these real-time tests were calculated to be 0.051, 0.060, and 0.121 mol% demonstrating the measurement systems’ capability to perform online monitoring with acceptable accuracy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

wtDigiTwin (Wind Turbine Digital Twin) [SWR-21-15]

This wind turbine digital twin software (wtDigiTwin) provides a digital twin solution for wind turbine applications. The focus of wtDigiTwin is to estimate loads, motions and environmental conditions for an operating wind turbine. The program uses supervisory control and data acquisition (SCADA) measurements as inputs, together with a wind turbine model. The wind industry is currently challenged by the high cost of operation and maintenance. These costs could be mitigated if component failures are predicted, but such predictions are difficult unless the turbines are equipped with expensive measuring devices. The alternative is to use a digital twin such as wtDigiTwin to estimate the necessary signals. wtDigiTwin can perform online prediction of signals that are otherwise not measured, using a limited set of reliable measurements and a physics-based model. The predicted signals can be used in applications that have direct cost benefits: 1) real-time estimation of the fatigue consumption of key components of the wind turbine; 2) root cause analyses and failure detections ; 3) lifetime reassessments ; 4) improvements to follow-on designs. The current version provides examples to estimate wind speed, thrust, torque, tower-top position, and tower loads on an onshore wind turbine using the following measurements tower top acceleration, generator torque, pitch, and rotational speed. The model combines a linear state-space model, a wind speed estimator, and a Kalman filter algorithm that integrates measurements with the state model to perform state estimations. The state space model is obtained either using OpenFAST linearizations, or using the yams package provided with the software.

Branlard, Emmanuel↗

CFM-ID 4.0 – a web server for accurate MS-based metabolite identification

The CFM-ID 4.0 web server (https://cfmid.wishartlab.com) is an online tool for predicting, annotating and interpreting tandem mass (MS/MS) spectra of small molecules. It is specifically designed to assist researchers pursuing studies in metabolomics, exposomics and analytical chemistry. More specifically, CFM-ID 4.0 supports the: 1) prediction of electrospray ionization quadrupole time-of-flight tandem mass spectra (ESI-QTOF-MS/MS) for small molecules over multiple collision energies (10 eV, 20 eV, and 40 eV); 2) annotation of ESI-QTOF-MS/MS spectra given the structure of the compound; and 3) identification of a small molecule that generated a given ESI-QTOF-MS/MS spectrum at one or more collision energies. The CFM-ID 4.0 web server makes use of a substantially improved MS fragmentation algorithm, a much larger database of experimental and in silico predicted MS/MS spectra and improved scoring methods to offer more accurate MS/MS spectral prediction and MS/MS-based compound identification. Compared to earlier versions of CFM-ID, this new version has an MS/MS spectral prediction performance that is ~22% better and a compound identification accuracy that is ~35% better on a standard (CASMI 2016) testing dataset. CFM-ID 4.0 also features a neutral loss function that allows users to identify similar or substituent compounds where no match can be found using CFM-ID’s regular MS/MS-to-compound identification utility. Finally, the CFM-ID 4.0 web server now offers a much more refined user interface that is easier to use, supports molecular formula identification (from MS/MS data), provides more interactively viewable data (including proposed fragment ion structures) and displays MS mirror plots for comparing predicted with observed MS/MS spectra. These improvements should make CFM-ID 4.0 much more useful to the community and should make small molecule identification much easier, faster, and more accurate.

59 BASIC BIOLOGICAL SCIENCES↗

Hybrid Imitation Learning for Real-Time Service Restoration in Resilient Distribution Systems

Self-healing capability is a critical factor for a resilient distribution system, which requires intelligent agents to automatically perform service restoration online, including network reconfiguration and reactive power dispatch. Here, the article proposes the imitation learning framework for training such an agent, where the agent will interact with an expert built based on the mixed-integer program to learn its optimal policy, and therefore significantly improve the training efficiency compared with exploration-dominant reinforcement learning (RL) methods. This significantly improved training efficiency makes the training problem under N-k scenarios tractable. A hybrid policy network is proposed to handle tie-line operations and reactive power dispatch simultaneously to further improve the restoration performance. The 33-bus and 119-bus systems with N-k disturbances are employed to conduct the training. The results indicate that the proposed method outperforms traditional RL algorithms such as the deep-Q network.

42 ENGINEERING↗

Three-dimensional modeling of hyphal fusion, branching, and nutrient transport in filamentous fungi

Fungi exhibit behaviors distinct from other microbes. Filamentous fungi grow by extending complex networks of branched filaments collectively referred to as the mycelium. These networks can expand over large distances and traverse low-nutrient areas by translocating nutrients through the filament network. This spatial characteristic makes filamentous fungi crucial for soil ecosystems, supporting stable microbial communities and promoting plant growth. However, simulating these behaviors is complex. The elongated nature of fungal compartments results in different mechanical interactions compared to the commonly modeled spherical bacteria. These detailed hyphal mechanics require specialized consideration and are often excluded from conventional fungal simulation packages. Additionally, the extensive fungal networks in nature demand computationally intensive simulations, necessitating high-performance algorithms. Therefore, realistic fungi simulations require specialized software. Here, we introduce a fungal modeling expansion to the high-performance biological modelling and interface exchange (bmx) software suite. bmx leverages adaptive mesh refinement in AMReX for chemical diffusion and incorporates a full mechanical model for bacterial cells, accelerated by GPUs. By extending bmx to model filamentous particles, we demonstrate the formation of complex filament networks through interactions like hyphal branching and fusion (anastomosis). We show that the networks produced match real-world fungal structures through various metrics. This work supports computational studies of fungal growth dynamics and can be adapted to investigate the growth of other filamentous structures in biology or materials science. The expanded-BMX package is open-sourced and is available online.

Cell mechanics↗

Kernelized approaches to streaming compression of scientific data

In this paper three algorithms are developed for the streaming compression of scientific data. The algorithms presented are reliant on the theory of vector-valued reproducing kernel Hilbert spaces and operator valued kernel. Further, the scientific data is modeled as a snapshot of time dependent vector field F(x, t) over a manifold M and the recovery of the data is framed as a learning problem. These processes are then appropriately modified and ana lyzed for the streaming scenario in which data is generated without the ability to revisit past entries.

97 MATHEMATICS AND COMPUTING↗

RAP: Resource-aware Automated GPU Sharing for Multi-GPU Recommendation Model Training and Input Preprocessing

Ensuring high-quality recommendations for newly onboarded users requires the continuous retraining of Deep Learning Recommendation Models (DLRMs) with freshly generated data. To serve the online DLRM retraining, existing solutions use hundreds of CPU computing nodes designated for input preprocessing, causing significant power consumption that surpasses even the power usage of GPU trainers. To this end, we propose RAP, an end-to-end DLRM training framework that supports Resource-aware Automated GPU sharing for DLRM input Preprocessing and Training. The core idea of RAP is to accurately capture the remaining GPU computing resources during DLRM training for input preprocessing, achieving superior training efficiency without requiring additional resources. Specifically, RAP utilizes a co-running cost model to efficiently assess the costs of various input preprocessing operations, and it implements a resource-aware horizontal fusion technique that adaptively merges smaller kernels according to GPU availability, circumventing any interference with DLRM training. In addition, RAP leverages a heuristic searching algorithm that jointly optimizes both the input preprocessing graph mapping and the co-running schedule to maximize the end-to-end DLRM training throughput. The comprehensive evaluation shows that RAP achieves 78.3× speedup on average over CPU-based DLRM input preprocessing frameworks. In addition, the end-to-end training throughput of RAP is only 2.04% lower than the ideal case, which has no input preprocessing overhead.

Wang, Zheng↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗