Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “pattern classification”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Systematic quark/gluon identification with ratios of likelihoods

Discriminating between quark- and gluon-initiated jets has long been a central focus of jet substructure, leading to the introduction of numerous observables and calculations to high perturbative accuracy. At the same time, there have been many attempts to fully exploit the jet radiation pattern using tools from statistics and machine learning. We propose a new approach that combines a deep analytic understanding of jet substructure with the optimality promised by machine learning and statistics. After specifying an approximation to the full emission phase space, we show how to construct the optimal observable for a given classification task. This procedure is demonstrated for the case of quark and gluons jets, where we show how to systematically capture sub-eikonal corrections in the splitting functions, and prove that linear combinations of weighted multiplicity is the optimal observable. In addition to providing a new and powerful framework for systematically improving jet substructure observables, we demonstrate the performance of several quark versus gluon jet tagging observables in parton-level Monte Carlo simulations, and find that they perform at or near the level of a deep neural network classifier. Combined with the rapid recent progress in the development of higher order parton showers, we believe that our approach provides a basis for systematically exploiting subleading effects in jet substructure analyses at the Large Hadron Collider (LHC) and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A spatial-statistical investigation of surface expressions associated with cyclic steaming in the Midway-Sunset Oil Field, California

In the Midway Sunset Oil Field in Central California, operators inject steam into the shallow diatomite formation to enhance heavy oil recovery through imbibition, wettability alteration, and viscosity reduction, among other mechanisms. The injected steam, however, does not always remain in the reservoir or return through the wells. In two zones in the study area, the steam comes out at the surface, creating sinkholes, seeps, and steam outlets. These phenomena, called “surface expressions,” pose safety and environmental hazards. Even though these surface expressions are a widespread problem in Central California, they are not well documented and understood. Possible causes of the surface expressions include: high injection pressure, structurally controlled flow patterns, leakage of steam through old improperly abandoned wells, high injection volumes, or flow along naturally occurring faults, among other possible factors. This work examines attributes of the zones with surface expressions in order to determine factors that may contribute to their occurrence. Spatial statistical analysis using logistic regression, random forests, and classification trees is used to explore the relationship between the surface expressions and geological and production-related attributes. The results point to a significant spatial correlation between the surface expressions and two predictors: concentration of plugged wells and geologic seal thickness. The results guide follow-up studies to further investigate the role of well abandonment and seal thickness in the occurrence of surface expressions.

02 PETROLEUM↗

Automated remote sensing tools to counter illicit maritime activity: vessel detection, bathymetry and topography from WorldView imagery

Illicit maritime activity, such as piracy and smuggling, is a global issue that often occurs in areas where interdiction authorities are sparse, ground-based monitoring technologies like radar lack range and coverage, and dark vessels abound. Our objective was to assess vessel congregation patterns, identify likely beaching areas based on terrain maps, and use them to narrow the search for potential smuggling transfer and overland routes in a specific region along the Puntland coast of Somalia. To accomplish this goal, we developed automated protocols applied to WorldView satellite imagery for (1) vessel detection and size classification, (2) shallow-water bathymetric characterization, and (3) coastal topography mapping. Utilizing a single sensor for vessel detection and topographic and bathymetric extraction at high spatial resolution and at a high re-visit rate provides a simplification to near-shore characterization for monitoring purposes. The extracted topography and bathymetry are then presented as a single, comprehensive perspective of a coastal region. The vessel-detection algorithm identified all vessels larger than approximately 15 metres in length (35 of 35), but misidentified six artefacts (i.e. false positives), resulting in an overall accuracy of 85%. The combined vessel and terrain maps facilitated the identification of a potential beaching and overland transportation route location.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Classification of Photovoltaic Failures with Hidden Markov Modeling, an Unsupervised Statistical Approach

Failure detection methods are of significant interest for photovoltaic (PV) site operators to help reduce gaps between expected and observed energy generation. Current approaches for field-based fault detection, however, rely on multiple data inputs and can suffer from interpretability issues. In contrast, this work offers an unsupervised statistical approach that leverages hidden Markov models (HMM) to identify failures occurring at PV sites. Using performance index data from 104 sites across the United States, individual PV-HMM models are trained and evaluated for failure detection and transition probabilities. This analysis indicates that the trained PV-HMM models have the highest probability of remaining in their current state (87.1% to 93.5%), whereas the transition probability from normal to failure (6.5%) is lower than the transition from failure to normal (12.9%) states. A comparison of these patterns using both threshold levels and operations and maintenance (O&M) tickets indicate high precision rates of PV-HMMs (median = 82.4%) across all of the sites. Although additional work is needed to assess sensitivities, the PV-HMM methodology demonstrates significant potential for real-time failure detection as well as extensions into predictive maintenance capabilities for PV.

classification↗

Learning Latent Interactions for Event Identification via Graph Neural Networks and PMU Data

Phasor measurement units (PMUs) are being widely installed on power systems, providing a unique opportunity to enhance wide-area situational awareness. One essential application is the use of PMU data for real-time event identification. However, how to take full advantage of all PMU data in event identification is still an open problem. Thus, we propose a novel method that performs event identification by mining interaction graphs among different PMUs. The proposed interaction graph inference method follows an entirely data-driven manner without knowing the physical topology. Moreover, unlike previous works that treat interactive learning and event identification as two different stages, our method learns interactions jointly with the identification task, thereby improving the accuracy of graph learning and ensuring seamless integration between the two stages. Moreover, to capture multi-scale event patterns, a dilated inception-based method is investigated to perform feature extraction of PMU data. To test the proposed data-driven approach, a large real-world dataset from tens of PMU sources and the corresponding event logs have been utilized in this work. We report numerical results validate that our method has higher classification accuracy compared to previous methods.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Two-Frequency RF Fields Induced Multipactor in Coaxial Transmission Lines

Multipactor is a nonlinear discharge phenomenon that occurs in vacuum RF systems, potentially leading to signal distortion, power loss, and even permanent damage to high-power components. This study presents a detailed investigation of two-surface multipactor in coaxial transmission lines under two-frequency excitation, using one-dimensional (1D) Monte Carlo simulations validated by three-dimensional (3D) Particle-in-Cell (PIC) results and experimental data. Introducing a second carrier mode is shown to suppress multipactor by reshaping and shrinking the susceptibility region, with the extent and location of suppression strongly dependent on the device aspect ratio and the relative phase of the second mode. Distinct suppression patterns emerge across different frequency–gap distance (fd) regimes, and in certain cases, susceptibility expansion is also observed. The study identifies and distinguishes pure and mixed multipactor modes in coaxial geometry, where analytical mode boundaries are not readily defined. Unlike planar systems, pure-mode regions in coaxial structures overlap with mixed-mode domains, complicating classification. Image charge forces are found to have minimal effect on susceptibility thresholds but do influence electron growth rates. These findings provide new insights into waveform-driven control of multipactor in high-power RF systems.

43 PARTICLE ACCELERATORS↗

A spatiotemporally explicit and scalable indicator of intact lands across the conterminous United States, 1986–2023

Globally, ecologically intact areas are increasingly scarce. Agricultural expansion into previously uncultivated areas drives the loss of intact lands that might otherwise exhibit high levels of ecological integrity. Thus, the absence of cultivation can be an indicator of intact lands as measured from remote sensing data and thematic maps. Our objective for this study was to develop and compare tractable approaches based on remotely sensed satellite data to map spatial patterns of potentially intact lands across the conterminous U.S. (CONUS). Using annual cultivation probabilities derived from satellite observations, we classified and mapped potentially intact lands across CONUS from 1986 to 2023 at 30 m resolution. We created three maps, first by applying a constant cultivation probability threshold across CONUS, second by varying the threshold state-by-state to maximize state-level overall accuracies, and third by equalizing the state-level user's and producer's accuracies to minimize classification bias. Validation against 800,000+ independent ground samples resulted in CONUS-level overall accuracies ≥85% for the roughly 660 million ha of potentially intact land. Map accuracy varied with the proportion of potentially intact lands across regions, with the Pacific-Mountain and Great Plains regions exhibiting the highest accuracies, while Eastern CONUS exhibited a greater mix of potentially intact and non-intact lands and more moderate map accuracies. These novel maps and approaches can be adapted to different spatiotemporal extents to support conservation and production decisions ranging from species and ecosystems protection to reducing land conversion and climate mitigation.

agriculture↗

Hierarchical Data-Driven Protection for Microgrid with 100% Renewable Penetration: Preprint

The accurate detection and isolation of faults is critical for the reliable operation of microgrids (MGs). Traditional protection approaches are even more challenged for 100% renewable MGs because inverter-based resources (IBRs) are the only sources for fault current which are usually low and unpredictable/non-uniform. This calls for new protection scheme that can identify IBR fault responses and detect faults in MGs. Data-driven based protection can learn the pattern of IBR fault responses and make the correct decision to identify faults. Therefore, this paper presents a data-driven approach for fault localization in island MGs. The approach builds a training dataset of comprehensive fault scenarios that can be used to learn fault characteristics from processed measurements. The localization task is modeled as a binary classification problem at each relay, which simplifies the learning process. Then, a hierarchical decision mechanism is used to identify the fault location. The proposed approach is assessed using an exemplary MG with several grid-forming (GFM) and grid-following (GFL) inverters, where accurate estimation of fault location is achieved. The data-driven based protection approach developed in this paper provides a generic framework and useful guidance for power system protection engineers to achieve reliable protection for MGs with 100% renewables.

artificial intelligence↗

Tetranucleotide frequencies differentiate genomic boundaries and metabolic strategies across environmental microbiomes

Microbiomes are constrained by physicochemical conditions, nutrient regimes, and community interactions across diverse environments, yet genomic signatures of this adaptation remain unclear. Metagenome sequencing is a powerful technique to analyze genomic content in the context of natural environments, establishing concepts of microbial ecological trends. Here, we developed a data discovery tool-a tetranucleotide-informed metagenome stability diagram-that is publicly available in the integrated microbial genomes and microbiomes (IMG/M) platform for metagenome ecosystem analyses. We analyzed the tetranucleotide frequencies from quality-filtered and unassembled sequence data of over 12,000 metagenomes to assess ecosystem-specific microbial community composition and function. We found that tetranucleotide frequencies can differentiate communities across various natural environments and that specific functional and metabolic trends can be observed in this structuring. Our tool places metagenomes sampled from diverse environments into clusters and along gradients of tetranucleotide frequency similarity, suggesting microbiome community compositions specific to gradient conditions. Within the resulting metagenome clusters, we identify protein-coding gene identifiers that are most differentiated between ecosystem classifications. We plan for annual updates to the metagenome stability diagram in IMG/M with new data, allowing for refinement of the ecosystem classifications delineated here. This framework has the potential to inform future studies on microbiome engineering, bioremediation, and the prediction of microbial community responses to environmental change. IMPORTANCE: Microbes adapt to diverse environments influenced by factors like temperature, acidity, and nutrient availability. We developed a new tool to analyze and visualize the genetic makeup of over 12,000 microbial communities, revealing patterns linked to specific functions and metabolic processes. This tool groups similar microbial communities and identifies characteristic genes within environments. By continually updating this tool, we aim to advance our understanding of microbial ecology, enabling applications like microbial engineering, bioremediation, and predicting responses to environmental change.

Kellom, Matthew↗

Uncovering acoustic signatures of pore formation in laser powder bed fusion

Abstract We present a machine learning workflow to discover signatures in acoustic measurements that can be utilized to create a low-dimensional model to accurately predict the location of keyhole pores formed during additive manufacturing processes. Acoustic measurements were sampled at 100 kHz during single-layer laser powder bed fusion (LPBF) experiments, and spatio-temporal registration of pore locations was obtained from post-build radiography. Power spectral density (PSD) estimates of the acoustic data were then decomposed using non-negative matrix factorization with custom $$\varvec{k}$$ k -means clustering (NMF $$\varvec{k}$$ k ) to learn the underlying spectral patterns associated with pore formation. NMF $$\varvec{k}$$ k returned a library of basis signals and matching coefficients to blindly construct a feature space based on the PSD estimates in an optimized fashion. Moreover, the NMF $$\varvec{k}$$ k decomposition led to the development of computationally inexpensive machine learning models which are capable of quickly and accurately identifying pore formation with classification accuracy of supervised and unsupervised label learning greater than 95% and 90%, respectively. The intrinsic data compression of NMF k , the relatively light computational cost of the machine learning workflow, and the high classification accuracy makes the proposed workflow an attractive candidate for edge computing toward in-situ keyhole pore prediction in LPBF.

36 MATERIALS SCIENCE↗

Analysis of biofilm assembly by large area automated AFM

Biofilms are complex microbial communities critical in medical, industrial, and environmental contexts. Understanding their assembly, structure, genetic regulation, interspecies interactions, and environmental responses is key to developing effective control and mitigation strategies. While atomic force microscopy (AFM) offers critically important high-resolution insights on structural and functional properties at the cellular and even sub-cellular level, its limited scan range and labor-intensive nature restricts the ability to link these smaller scale features to the functional macroscale organization of the films. We begin to address this limitation by introducing an automated large area AFM approach capable of capturing high-resolution images over millimeter-scale areas, aided by machine learning for seamless image stitching, cell detection, and classification. Large area AFM is shown to provide a very detailed view of spatial heterogeneity and cellular morphology during the early stages of biofilm formation which were previously obscured. Using this approach, we examined the organization of Pantoea sp. YR343 on PFOTS-treated glass surfaces. Our findings reveal a preferred cellular orientation among surface-attached cells, forming a distinctive honeycomb pattern. Detailed mapping of flagella interactions suggests that flagellar coordination plays a role in biofilm assembly beyond initial attachment. Additionally, we use large-area AFM to characterize surface modifications on silicon substrates, observing a significant reduction in bacterial density. This highlights the potential of this method for studying surface modifications to better understand and control bacterial adhesion and biofilm formation.

59 BASIC BIOLOGICAL SCIENCES↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

Disentangling Alzheimer’s disease neurodegeneration from typical brain ageing using machine learning

Abstract Neuroimaging biomarkers that distinguish between changes due to typical brain ageing and Alzheimer’s disease are valuable for determining how much each contributes to cognitive decline. Supervised machine learning models can derive multivariate patterns of brain change related to the two processes, including the Spatial Patterns of Atrophy for Recognition of Alzheimer’s Disease (SPARE-AD) and of Brain Aging (SPARE-BA) scores investigated herein. However, the substantial overlap between brain regions affected in the two processes confounds measuring them independently. We present a methodology, and associated results, towards disentangling the two. T1-weighted MRI scans of 4054 participants (48–95 years) with Alzheimer’s disease, mild cognitive impairment (MCI), or cognitively normal (CN) diagnoses from the Imaging-based coordinate SysTem for AGIng and NeurodeGenerative diseases (iSTAGING) consortium were analysed. Multiple sets of SPARE scores were investigated, in order to probe imaging signatures of certain clinically or molecularly defined sub-cohorts. First, a subset of clinical Alzheimer’s disease patients (n = 718) and age- and sex-matched CN adults (n = 718) were selected based purely on clinical diagnoses to train SPARE-BA1 (regression of age using CN individuals) and SPARE-AD1 (classification of CN versus Alzheimer’s disease) models. Second, analogous groups were selected based on clinical and molecular markers to train SPARE-BA2 and SPARE-AD2 models: amyloid-positive Alzheimer’s disease continuum group (n = 718; consisting of amyloid-positive Alzheimer’s disease, amyloid-positive MCI, amyloid- and tau-positive CN individuals) and amyloid-negative CN group (n = 718). Finally, the combined group of the Alzheimer’s disease continuum and amyloid-negative CN individuals was used to train SPARE-BA3 model, with the intention to estimate brain age regardless of Alzheimer’s disease-related brain changes. The disentangled SPARE models, SPARE-AD2 and SPARE-BA3, derived brain patterns that were more specific to the two types of brain changes. The correlation between the SPARE-BA Gap (SPARE-BA minus chronological age) and SPARE-AD was significantly reduced after the decoupling (r = 0.56–0.06). The correlation of disentangled SPARE-AD was non-inferior to amyloid- and tau-related measurements and to the number of APOE ε4 alleles but was lower to Alzheimer’s disease-related psychometric test scores, suggesting the contribution of advanced brain ageing to the latter. The disentangled SPARE-BA was consistently less correlated with Alzheimer’s disease-related clinical, molecular and genetic variables. By employing conservative molecular diagnoses and introducing Alzheimer’s disease continuum cases to the SPARE-BA model training, we achieved more dissociable neuroanatomical biomarkers of typical brain ageing and Alzheimer’s disease.

Hwang, Gyujoon↗

A high-throughput workflow to analyze sequence-conformation relationships and explore hydrophobic patterning in disordered peptoids

Understanding how a macromolecule’s primary sequence governs its conformational landscape is crucial for elucidating its function, yet these design principles are still emerging for macromolecules with intrinsic disorder. Herein, we introduce a high-throughput workflow that implements a practical colorimetric conformational assay, introduces a semi-automated sequencing protocol using matrix-assisted laser desorption/ionization and tandem mass spectrometry (MALDI-MS/MS), and develops a generalizable sequence-structure algorithm. Using a model system of 20mer peptidomimetics containing polar glycine and hydrophobic N-butylglycine residues, we identified nine classifications of conformational disorder and isolated 122 unique sequences across varied compositions and conformations. Conformational distributions of three compositionally identical library sequences were corroborated through atomistic simulations and ion mobility spectrometry coupled with liquid chromatography. A data-driven strategy was developed using existing sequence variables and data-derived “motifs” to inform a machine-learning algorithm toward conformation prediction. Here, this multifaceted approach enhances our understanding of sequence-conformation relationships and offers a powerful tool for accelerating the discovery of materials with conformational control.

data-driven analysis↗

SigTime: Learning and Visually Explaining Time Series Signatures

Understanding and distinguishing temporal patterns in time series data is essential for scientific discovery and decision-making. For example, in biomedical research, uncovering meaningful patterns in physiological signals can improve diagnosis, risk assessment, and patient outcomes. However, existing methods for time series pattern discovery face major challenges, including high computational complexity, limited interpretability, and difficulty in capturing meaningful temporal structures. Here, to address these gaps, we introduce a novel learning framework that jointly trains two Transformer models using complementary time series representations: shapelet-based representations to capture localized temporal structures and traditional feature engineering to encode statistical properties. The learned shapelets serve as interpretable signatures that differentiate time series across classification labels. Additionally, we develop a visual analytics system—SigTime—with coordinated views to facilitate exploration of time series signatures from multiple perspectives, aiding in useful insights generation. We quantitatively evaluate our learning framework on eight publicly available datasets and one proprietary clinical dataset. Additionally, we demonstrate the effectiveness of our system through two usage scenarios along with the domain experts: one involving public ECG data and the other focused on preterm labor analysis.

97 MATHEMATICS AND COMPUTING↗

Composition and metabolism of microbial communities in soil pores

Delineation of microbial habitats within the soil matrix and characterization of their environments and metabolic processes are crucial to understand soil functioning, yet their experimental identification remains persistently limited. We combined single- and triple-energy X-ray computed microtomography with pore specific allocation of 13 C labeled glucose and subsequent stable isotope probing to demonstrate how long-term disparities in vegetation history modify spatial distribution patterns of soil pore and particulate organic matter drivers of microbial habitats, and to probe bacterial communities populating such habitats. Here we show striking differences between large (30-150 µm Ø) and small (4-10 µm Ø) soil pores in (i) microbial diversity, composition, and life-strategies, (ii) responses to added substrate, (iii) metabolic pathways, and (iv) the processing and fate of labile C. We propose a microbial habitat classification concept based on biogeochemical mechanisms and localization of soil processes and also suggests interventions to mitigate the environmental consequences of agricultural management.

59 BASIC BIOLOGICAL SCIENCES↗

The K2 Galactic Archaeology Program Data Release 3: Age-abundance Patterns in C1–C8 and C10–C18

We present the third and final data release of the K2 Galactic Archaeology Program (K2 GAP) for Campaigns C1–C8 and C10–C18. We provide asteroseismic radius and mass coefficients, κ R and κ M , for ~19,000 red giant stars, which translate directly to radius and mass given a temperature. As such, K2 GAP DR3 represents the largest asteroseismic sample in the literature to date. K2 GAP DR3 stellar parameters are calibrated to be on an absolute parallactic scale based on Gaia DR2, with red giant branch and red clump evolutionary state classifications provided via a machine-learning approach. Combining these stellar parameters with GALAH DR3 spectroscopy, we determine asteroseismic ages with precisions of ~20%–30% and compare age-abundance relations to Galactic chemical evolution models among both low- and high-α populations for α, light, iron-peak, and neutron-capture elements. We confirm recent indications in the literature of both increased Ba production at late Galactic times as well as significant contributions to r-process enrichment from prompt sources associated with, e.g., core-collapse supernovae. With an eye toward other Galactic archeology applications, we characterize K2 GAP DR3 uncertainties and completeness using injection tests, suggesting that K2 GAP DR3 is largely unbiased in mass/age, with uncertainties of 2.9% (stat.) ± 0.1% (syst.) and 6.7% (stat.) ± 0.3% (syst.) in κ R and κ M for red giant branch stars and 4.7% (stat.) ± 0.3% (syst.) and 11% (stat.) ± 0.9% (syst.) for red clump stars. We also identify percent-level asteroseismic systematics, which are likely related to the time baseline of the underlying data, and which therefore should be considered in TESS asteroseismic analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Predicting Large‐Scale Systematic Missing Pipe Attributes in Water Distribution Networks

Water distribution network (WDN) models are an essential tool used by water utilities for hydraulic analysis. Unfortunately, missing data and insufficient resources often make creating and maintaining these models unfeasible. Existing methods to address missing pipe properties, like sequential imputation for missing values and reconstruction using graph metrics, are designed to accommodate random patterns of missing information and require a significant percentage of the system's attributes to be known. However, these data completeness assumptions do not always align with real‐world scenarios where large sections of the WDN model have missing data. To address this challenge, this study proposes a data‐driven approach for estimating pipe diameter when considering different spatial patterns and degrees of data completeness (i.e., 0%–90%). Using data from 16 WDNs in Kentucky, this study compares the use of machine learning (ML) using topological and geospatial features against an existing deterministic approach. Results demonstrate that WDN models with pipe diameters predicted by the proposed ML method had comparable hydraulic performance to the ground truth models. Moreover, results showed that ML method performance varies between WDNs of differing topological classification. Insights from this study help advance the ability to leverage partial data to create and maintain WDN models amid uncertainty and inadequate resources.

Poff, Jason W. [Oregon State Univ., Corvallis, OR ↗