Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “read classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Metagenome-assembled genomes from Wind River Basin floodplain sediments Riverton, Wyoming site (May to September 2017)

Microorganisms play a key role in cycling nutrients and contaminants in the terrestrial environment depending on their genetic potential. Here we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community in floodplain sediment samples taken roughly every month in the period May 18 to September 13 in 2017 at a location (Pit2) close to DOE Legacy Management well 855 at the Riverton, Wyoming floodplain site in the Wind River Basin (WRB). The groundwater at this site exhibits persistent U, Mo, and sulfate plumes and is one of the field sites in focus for the SLAC Groundwater Quality SFA program. Cores were taken with a hand-auger and separated into 5-20 cm segments based on soil horizonation down to 150 cm depth below surface. Each segment was subsampled for microbial analyses. Corresponding 16S rRNA gene amplicon data is available at the NCBI Single Read Archive (SRA) Database BioProject ID PRJNA626616, and soil geochemistry data at doi:10.15485/1631972. 40 metagenomes were sequenced through JGI and can be found under Gold sequencing project: Gs0142591. Metagenomes were assembled, binned, and refined using metawrap to generate MAGs (>50% complete and < 10% contamination based on checkM scores). This dataset includes a zip file of 6993 MAG fasta files and a csv file with quality, taxonomic classification (GTDB RS220), and metagenome accessions for MAGs generated from the Wind River Basin (WRB). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Chemical Sensing of Unexploded Ordnance with the Mobile Underwater Survey System (MUDSS)

The ability to sense explosives residues in the marine environment is a critical tool for identification and classification of underwater unexploded ordnance (UXO). Trace explosives signatures of TNT and DNT have been extracted from multiple sediment samples adjacent to unexploded undersea ordnance at Halifax Harbor, Canada. The ordnance was hurled into the harbor during a massive explosion fifty years earlier, in 1945 after World War II had ended. Laboratory sediment extractions were made using the solid-phase microextraction (SPME) method in seawater, and detection using the Reversal Electron Attachment Detection (READ) technique and, in the case of DNT, a commercial gas-chromatography/mass spectrometer (GC/MS). Results show that, after more than 50 years in the environment, ordnance which appeared to be physically intact gave good explosives signatures at the parts-per-billion level, whereas ordnance which had been cracked open during the explosion gave no signatures at the 10 parts-per-trillion sensitivity level. These measurements appear to provide the first reported data of explosives signatures from undersea UXOs.

Darrach, M. R.↗

Signal detection theory and methods for evaluating human performance in decision tasks

Signal Detection Theory (SDT) can be used to assess decision making performance in tasks that are not commonly thought of as perceptual. SDT takes into account both the sensitivity and biases in responding when explaining the detection of external events. In the standard SDT tasks, stimuli are selected in order to reveal the sensory capabilities of the observer. SDT can also be used to describe performance when decisions must be made as to the classification of easily and reliably sensed stimuli. Numbers are stimuli that are minimally affected by sensory processing and can belong to meaningful categories that overlap. Multiple studies have shown that the task of categorizing numbers from overlapping normal distributions produces performance predictable by SDT. These findings are particularly interesting in view of the similarity between the task of the categorizing numbers and that of determining the status of a mechanical system based on numerical values that represent sensor readings. Examples of the use of SDT to evaluate performance in decision tasks are reviewed. The methods and assumptions of SDT are shown to be effective in the measurement, evaluation, and prediction of human performance in such tasks.

Obrien, Kevin↗

Automatic information extraction from childhood cancer pathology reports

The International Classification of Childhood Cancer (ICCC) facilitates the effective classification of a heterogeneous group of cancers in the important pediatric population. However, there has been no development of machine learning models for the ICCC classification. We developed deep learning-based information extraction models from cancer pathology reports based on the ICD-O-3 coding standard. In this article, we describe extending the models to perform ICCC classification. We developed 2 models, ICD-O-3 classification and ICCC recoding (Model 1) and direct ICCC classification (Model 2), and 4 scenarios subject to the training sample size. We evaluated these models with a corpus consisting of 29206 reports with age at diagnosis between 0 and 19 from 6 state cancer registries. Our findings suggest that the direct ICCC classification (Model 2) is substantially better than reusing the ICD-O-3 classification model (Model 1). Applying the uncertainty quantification mechanism to assess the confidence of the algorithm in assigning a code demonstrated that the model achieved a micro-F1 score of 0.987 while abstaining (not sufficiently confident to assign a code) on only 14.8% of ambiguous pathology reports. Our experimental results suggest that the machine learning-based automatic information extraction from childhood cancer pathology reports in the ICCC is a reliable means of supplementing human annotators at state cancer registries by reading and abstracting the majority of the childhood cancer pathology reports accurately and reliably.

60 APPLIED LIFE SCIENCES↗

A starting guide to root ecology: strengthening ecological concepts and standardising root classification, sampling, processing and trait measurements

In the context of a recent massive increase in research on plant root functions and their impact on the environment, root ecologists currently face many important challenges to keep on generating cutting-edge, meaningful and integrated knowledge. Consideration of the below-ground components in plant and ecosystem studies has been consistently called for in recent decades, but methodology is disparate and sometimes inappropriate. This handbook, based on the collective effort of a large team of experts, will improve trait comparisons across studies and integration of information across databases by providing standardised methods and controlled vocabularies. It is meant to be used not only as starting point by students and scientists who desire working on below-ground ecosystems, but also by experts for consolidating and broadening their views on multiple aspects of root ecology. Beyond the classical compilation of measurement protocols, we have synthesised recommendations from the literature to provide key background knowledge useful for: (1) defining below-ground plant entities and giving keys for their meaningful dissection, classification and naming beyond the classical fine-root vs coarse-root approach; (2) considering the specificity of root research to produce sound laboratory and field data; (3) describing typical, but overlooked steps for studying roots (e.g. root handling, cleaning and storage); and (4) gathering metadata necessary for the interpretation of results and their reuse. Most importantly, all root traits have been introduced with some degree of ecological context that will be a foundation for understanding their ecological meaning, their typical use and uncertainties, and some methodological and conceptual perspectives for future research. Considering all of this, we urge readers not to solely extract protocol recommendations for trait measurements from this work, but to take a moment to read and reflect on the extensive information contained in this broader guide to root ecology, including sections I–VII and the many introductions to each section and root trait description. Finally, it is critical to understand that a major aim of this guide is to help break down barriers between the many subdisciplines of root ecology and ecophysiology, broaden researchers’ views on the multiple aspects of root study and create favourable conditions for the inception of comprehensive experiments on the role of roots in plant and ecosystem functioning.

59 BASIC BIOLOGICAL SCIENCES↗

Monitoring water quality from LANDSAT

Water quality monitoring possibilities from LANDSAT were demonstrated both for direct readings of reflectances from the water and indirect monitoring of changes in use of land surrounding Swift Creek Reservoir in a joint project with the Virginia State Water Control Board and NASA. Film products were shown to have insufficient resolution and all work was done by digitally processing computer compatible tapes. Land cover maps of the 18,000 hectare Swift Creek Reservoir watershed, prepared for two dates in 1974, are shown. A significant decrease in the pine cover was observed in a 740 hectare construction site within the watershed. A measure of the accuracy of classification was obtained by comparing the LANDSAT results with visual classification at five sites on a U-2 photograph. Such changes in land cover can alert personnel to watch for potential changes in water quality.

Barker, J. L.↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

A Step Beyond Simple Keyword Searches: Services Enabled by a Full Content Digital Journal Archive

The problems of managing and searching large archives of scientific journal articles can potentially be addressed through data mining and statistical techniques matured primarily for quantitative scientific data analysis. A journal paper could be represented by a multivariate descriptor, e.g., the occurrence counts of a number key technical terms or phrases (keywords), perhaps derived from a controlled vocabulary ( e . g . , the American Meteorological Society's Glossary of Meteorology) or bootstrapped from the journal archive itself. With this technique, conventional statistical classification tools can be leveraged to address challenges faced by both scientists and professional societies in knowledge management. For example, cluster analyses can be used to find bundles of "most-related" papers, and address the issue of journal bifurcation (when is a new journal necessary, and what topics should it encompass). Similarly, neural networks can be trained to predict the optimal journal (within a society's collection) in which a newly submitted paper should be published. Comparable techniques could enable very powerful end-user tools for journal searches, all premised on the view of a paper as a data point in a multidimensional descriptor space, e.g.: "find papers most similar to the one I am reading", "build a personalized subscription service, based on the content of the papers I am interested in, rather than preselected keywords", "find suitable reviewers, based on the content of their own published works", etc. Such services may represent the next "quantum leap" beyond the rudimentary search interfaces currently provided to end-users, as well as a compelling value-added component needed to bridge the print-to-digital-medium gap, and help stabilize professional societies' revenue stream during the print-to-digital transition.

Boccippio, Dennis J.↗

Seedling nitrogen uptake and rhizodeposition between mycorrhizal types

Tree mycorrhizal associations are associated with patterns in N cycling and soil organic matter (SOM) storage, however, we still lack a mechanistic understanding of what tree and fungal traits drive these patterns and how they will respond to global changes in soil N availability. To address this knowledge gap, we investigated how arbuscular mycorrhizal (AM)- and ectomycorrhizal (EcM)-associated seedlings alter rhizodeposition in response to increased inorganic N acquisition. Specifically, we conducted this greenhouse experiment in a sealed labeling chamber with an enriched 13Carbon atmosphere and 15Nitrogen enriched fertilizer over the course of five months from April 2021 - August 2021. To include the variability across tree species, we grew eight species of seedlings belonging to eight families that were either arbuscular (Acer rubrum, Nyssa sylvatica, Thuja occidentalis, and Prunus seritina) or ectomycorrhizal-associated (Quercus rubra, Tilia americana, Pinus strobus, and Betula lenta). We measured rhizodeposition (mg 13C), plant N uptake from fertilizer (mg N), net soil carbon, and the abundance of mycorrhizal fungi (ITS sequencing and qPCR). We also characterized fungal (ITS2) and bacterial (16S) soil communities.The data from this project are ".csv" files that can up downloaded into a folder, and then run in the associated R markdown scripts after changing the source folder location at the top of the script. These data include raw outputs and processed files (using the R markdown files) for seedling growth and biomass, 15N content, soil 13C content, raw reads and processed file versions for fungal and bacterial ASVS, and a final summary file used for modeling. R software is needed to run these data, and the packages needed are listed at the top of the R markdown file.

54 ENVIRONMENTAL SCIENCES↗

Automated Fiber Placement Defects: Automated Inspection and Characterization

Automated Fiber Placement (AFP) is an additive composite manufacturing technique, and a pressing challenge facing this technology is defect detection and repair. Manual defect inspection is time consuming, which led to the motivation to develop a rapid automatic method of inspection. This paper suggests a new automated inspection system based on convolutional neural networks and image segmentation tasks. This creates a pixel by pixel classification of the defects of the whole part scan. This process will allow for greater defect information extraction and faster processing times over previous systems, motivating rapid part inspection and analysis. Fine shape, height, and boundary detail can be generated through our system as opposed to a more coarse resolution demonstrated in other techniques. These scans are analyzed for defects, and then each defect is stored for export, or correlated to machine parameters or part design. The network is further improved through novel optimization techniques. New training instances can also be created with every new part scan by including the machine operator as a post inspection check on the accuracy of the system. Having a continuously adapting inspection system will increase accuracy for automated inspections, cutting down on false readings.

Sacco, Christopher↗

Visual information processing II; Proceedings of the Meeting, Orlando, FL, Apr. 14-16, 1993

Various papers on visual information processing are presented. Individual topics addressed include: aliasing as noise, satellite image processing using a hammering neural network, edge-detetion method using visual perception, adaptive vector median filters, design of a reading test for low-vision image warping, spatial transformation architectures, automatic image-enhancement method, redundancy reduction in image coding, lossless gray-scale image compression by predictive GDF, information efficiency in visual communication, optimizing JPEG quantization matrices for different applications, use of forward error correction to maintain image fidelity, effect of peanoscanning on image compression. Also discussed are: computer vision for autonomous robotics in space, optical processor for zero-crossing edge detection, fractal-based image edge detection, simulation of the neon spreading effect by bandpass filtering, wavelet transform (WT) on parallel SIMD architectures, nonseparable 2D wavelet image representation, adaptive image halftoning based on WT, wavelet analysis of global warming, use of the WT for signal detection, perfect reconstruction two-channel rational filter banks, N-wavelet coding for pattern classification, simulation of image of natural objects, number-theoretic coding for iconic systems.

Huck, Friedrich O.↗

NANO.PTML model for read-across prediction of nanosystems in neurosciences. computational model and experimental case of study

Abstract Neurodegenerative diseases involve progressive neuronal death. Traditional treatments often struggle due to solubility, bioavailability, and crossing the Blood-Brain Barrier (BBB). Nanoparticles (NPs) in biomedical field are garnering growing attention as neurodegenerative disease drugs (NDDs) carrier to the central nervous system. Here, we introduced computational and experimental analysis. In the computational study, a specific IFPTML technique was used, which combined Information Fusion (IF) + Perturbation Theory (PT) + Machine Learning (ML) to select the most promising Nanoparticle Neuronal Disease Drug Delivery (N2D3) systems. For the application of IFPTML model in the nanoscience, NANO.PTML is used. IF-process was carried out between 4403 NDDs assays and 260 cytotoxicity NP assays conducting a dataset of 500,000 cases. The optimal IFPTML was the Decision Tree (DT) algorithm which shown satisfactory performance with specificity values of 96.4% and 96.2%, and sensitivity values of 79.3% and 75.7% in the training (375k/75%) and validation (125k/25%) set. Moreover, the DT model obtained Area Under Receiver Operating Characteristic (AUROC) scores of 0.97 and 0.96 in the training and validation series, highlighting its effectiveness in classification tasks. In the experimental part, two samples of NPs (Fe 3 O 4 _A and Fe 3 O 4 _B) were synthesized by thermal decomposition of an iron(III) oleate (FeOl) precursor and structurally characterized by different methods. Additionally, in order to make the as-synthesized hydrophobic NPs (Fe 3 O 4 _A and Fe 3 O 4 _B) soluble in water the amphiphilic CTAB (Cetyl Trimethyl Ammonium Bromide) molecule was employed. Therefore, to conduct a study with a wider range of NP system variants, an experimental illustrative simulation experiment was performed using the IFPTML-DT model. For this, a set of 500,000 prediction dataset was created. The outcome of this experiment highlighted certain NANO.PTML systems as promising candidates for further investigation. The NANO.PTML approach holds potential to accelerate experimental investigations and offer initial insights into various NP and NDDs compounds, serving as an efficient alternative to time-consuming trial-and-error procedures.

60 APPLIED LIFE SCIENCES↗

Classification of bacterial plasmid and chromosome derived sequences using machine learning

Plasmids are important genetic elements that facilitate horizonal gene transfer between bacteria and contribute to the spread of virulence and antimicrobial resistance. Most bacterial genome sequences in the public archives exist in draft form with many contigs, making it difficult to determine if a contig is of chromosomal or plasmid origin. Using a training set of contigs comprising 10,584 chromosomes and 10,654 plasmids from the PATRIC database, we evaluated several machine learning models including random forest, logistic regression, XGBoost, and a neural network for their ability to classify chromosomal and plasmid sequences using nucleotide k-mers as features. Based on the methods tested, a neural network model that used nucleotide 6-mers as features that was trained on randomly selected chromosomal and plasmid subsequences 5kb in length achieved the best performance, outperforming existing out-of-the-box methods, with an average accuracy of 89.38% ± 2.16% over a 10-fold cross validation. The model accuracy can be improved to 92.08% by using a voting strategy when classifying holdout sequences. In both plasmids and chromosomes, subsequences encoding functions involved in horizontal gene transfer—including hypothetical proteins, transporters, phage, mobile elements, and CRISPR elements—were most likely to be misclassified by the model. This study provides a straightforward approach for identifying plasmid-encoding sequences in short read assemblies without the need for sequence alignment-based tools.

59 BASIC BIOLOGICAL SCIENCES↗

The adult literacy evaluator: An intelligent computer-aided training system for diagnosing adult illiterates

An important part of NASA's mission involves the secondary application of its technologies in the public and private sectors. One current application being developed is The Adult Literacy Evaluator, a simulation-based diagnostic tool designed to assess the operant literacy abilities of adults having difficulties in learning to read and write. Using ICAT system technology in addition to speech recognition, closed-captioned television (CCTV), live video and other state-of-the art graphics and storage capabilities, this project attempts to overcome the negative effects of adult literacy assessment by allowing the client to interact with an intelligent computer system which simulates real-life literacy activities and materials and which measures literacy performance in the actual context of its use. The specific objectives of the project are as follows: (1) To develop a simulation-based diagnostic tool to assess adults' prior knowledge about reading and writing processes in actual contexts of application; (2) to provide a profile of readers' strengths and weaknesses; and (3) to suggest instructional strategies and materials which can be used as a beginning point for remediation. In the first and developmental phase of the project, descriptions of literacy events and environments are being written and functional literacy documents analyzed for their components. Examples of literacy events and situations being considered included interactions with environmental print (e.g., billboards, street signs, commercial marquees, storefront logos, etc.), functional literacy materials (e.g., newspapers, magazines, telephone books, bills, receipts, etc.) and employment related communication (i.e., job descriptions, application forms, technical manuals, memorandums, newsletters, etc.). Each of these situations and materials is being analyzed for its literacy requirements in terms of written display (i.e., knowledge of printed forms and conventions), meaning demands (i.e., comprehension and word knowledge) and social situation. From these descriptions, scripts are being generated which define the interaction between the student, an on-screen guide and the simulated literacy environment. The proposed outcome of the Evaluator is a diagnostic profile which will present broad classifications of literacy behaviors across the major areas of metacognitive abilities, word recognition, vocabulary knowledge, comprehension and writing. From these classifications, suggestions for materials and strategies for instruction with which to begin corrective action will be made. The focus of the Literacy Evaluator will be essentially to provide an expert diagnosis and an interpretation of that assessment which then can be used by a human tutor to further design and individualize a remedial program as needed through the use of an authoring system.

Yaden, David B., Jr.↗

Classification of River Catchments in the Contiguous United States: Code, Dataset, Similarity Patterns, and Resulting Classes

This dataset serves as supplementary information for the paper by Ciulla F. and Varadharajan C. A Network Approach for Multiscale Catchment Classification using Traits (see reference 1). It contains environmental and physical catchment traits, such as temperatures, precipitation, land use and human interference, from 9067 sites across the contiguous United States (CONUS). The purpose of this dataset is to provide information for a better trait-based categorization of river catchments in the CONUS using networks as an analytical tool. The traits variables match the ones present in the GAGES-II dataset and the preprocessing steps are described in the Methods section (processed_dataset.csv). Additionally we include the topologies (nodes, edges and clusters, also referred as classes) of the catchment network and traits network generated by said dataset (csv and json files). A series of tables support the information carried by the network providing more detailed descriptions of cluster components (SI1.pdf). A summary of all the plots of clusters of catchments with at least 50 nodes is provided (SI2.pdf). The characteristic traits for each cluster of catchments is presented as z-score (traits_categories_zscores_per_catchment_class.csv). The link to the hydrological behavior of clusters of catchments is displayed by boxplots, each describing a particular river discharge index (SI3.pdf). Both csv and json files can be read by common text editors but the data contained into them can be better handled using programming languages like python and database oriented libraries like pandas. Pdf files can be read by any pdf reader software.[02-23-2024] Update: The code and datasets necessary to reproduce the results of the study are available as a zipped repository (code_datasets_catchments_similarity.zip).

54 ENVIRONMENTAL SCIENCES↗

Relative effectiveness of kinetic analysis vs single point readings for classifying environmental samples based on community-level physiological profiles (CLPP)

The relative effectiveness of average-well-color-development-normalized single-point absorbance readings (AWCD) vs the kinetic parameters mu(m), lambda, A, and integral (AREA) of the modified Gompertz equation fit to the color development curve resulting from reduction of a redox sensitive dye from microbial respiration of 95 separate sole carbon sources in microplate wells was compared for a dilution series of rhizosphere samples from hydroponically grown wheat and potato ranging in inoculum densities of 1 x 10(4)-4 x 10(6) cells ml-1. Patterns generated with each parameter were analyzed using principal component analysis (PCA) and discriminant function analysis (DFA) to test relative resolving power. Samples of equivalent cell density (undiluted samples) were correctly classified by rhizosphere type for all parameters based on DFA analysis of the first five PC scores. Analysis of undiluted and 1:4 diluted samples resulted in misclassification of at least two of the wheat samples for all parameters except the AWCD normalized (0.50 abs. units) data, and analysis of undiluted, 1:4, and 1:16 diluted samples resulted in misclassification for all parameter types. Ordination of samples along the first principal component (PC) was correlated to inoculum density in analyses performed on all of the kinetic parameters, but no such influence was seen for AWCD-derived results. The carbon sources responsible for classification differed among the variable types with the exception of AREA and A, which were strongly correlated. These results indicate that the use of kinetic parameters for pattern analysis in CLPP may provide some additional information, but only if the influence of inoculum density is carefully considered. c2001 Elsevier Science Ltd. All rights reserved.

NASA Center KSC↗

Depth-resolved sagebrush root metabolomics, rhizosphere microbial communities, and geochemistry at the East River Watershed

This data set consists of results from soil nutrient profile, untargeted metabolomics, mass spec imaging, and amplicon sequencing. Data for soil nutrient profile includes common cations (Ca, Mg, Na, and K etc.) extracted from 3 digesting steps – ammonia acetate (for exchangeable cations), nitric acid (for acid dissolved fraction), and hydrofluoric acid/perchloric acid (HF/HClO4) for whole soil digestion. It also includes concentration of organic carbon, inorganic nitrogen (ammonia and nitrate) and phosphorus (Bray-1 P and nitric acid extract), and total nitrogen and phosphorus. Data for untargeted metabolomics includes metabolomic profile for root exudate/tissues and soil extracts from depths at surface soil to saprolite, that were measured using gas chromatography – mass spectrometry (GC-MS), and liquid chromatography – tandem mass spectrometry (LC-MS/MS). Data for mass spec imaging includes spatial distribution of metabolites that were detected and annotated with Fourier transformation ion cyclotron resonance mass spectrometer (FTICR-MS). Data for amplicon sequencing includes the base paired 16S and ITS ribosomal RNA sequences from Miseq Illumina sequencing. All samples were collected from 2 sampling campaign October 2022 and June 2023. Collectively, these datasets enable a mechanistic evaluation of how nutrient acquisition, especially nitrogen and phosphorus, differs between shallow roots operating in soil and deep roots functioning within the fractured bedrock zone. All files are provided as comma-separated values (CSV) fies (.csv) and (GZIP) file (.gz). The compressed .gz FASTQ files can be read directly in R using the dada2 package as part of the amplicon sequence analysis workflow. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research was performed on a project award 60563 (https://dx.doi.org/10.46936/expl.proj.2022.60563/60008727) from the Environmental Molecular Sciences Laboratory, a DOE Office of Science User Facility sponsored by the Biological and Environmental Research program under Contract No. DE-AC05-76RL01830.

EARTH SCIENCE > AGRICULTURE > SOILS > CARBON↗

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN↗