Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Markov models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Oil price states and drivers: An analysis of the second-month spot-futures price differential

Oil remains a dominant component of global energy use, and its price, characterized by frequent changes and an ever-present potential for large swings, continues to be a focus of industry participants, policymakers and analysts attention. Here, this study examines the behavior of future spot oil prices using a continuous-time hidden Markov model (HMM) and daily price data spanning years 2007 to 2024. We identify six states in the second-month WTI spot-futures price differential and assess the roles of eleven futures price, quantity, financial market, and geopolitical/volatility variables in each state. The model yields several insights into the workings of the oil market and the relative roles of these drivers. We find support for several theoretical and empirical findings in the oil market literature, including the role of inventory, volatility/risk, and market responses to contango/backwardation in futures markets. A novel finding is that “normal contango” conditions represent a significant portion of next-day states in our in-sample data. Under the most volatile normal contango state, many of the oil market drivers differ markedly in coefficient signs and magnitudes from those in other states. The resulting model also performed well out-of-sample and would, in addition to understanding the impact of market drivers, be useful for short-term forecasting. Overall, the findings highlight the highly non-linear, regime-dependent interactions of the oil price and its drivers, emphasizing the importance of detailed information to market stakeholders.

Oladosu, Gbadebo A. [Oak Ridge National Laboratory↗

Improving Estimation of the Koopman Operator with Kolmogorov–Smirnov Indicator Functions

It has become common to perform kinetic analysis using approximate Koopman operators that transform high-dimensional timeseries of observables into ranked dynamical modes. The key to the practical success of the approach is the identification of a set of observables that form a good basis on which to expand the slow relaxation modes. Good observables are, however, difficult to identify a priori and suboptimal choices can lead to significant underestimations of characteristic time scales. Leveraging the representation of slow dynamics in terms of Hidden Markov Models (HMM), we propose a simple and computationally efficient clustering procedure to infer surrogate observables that form a good basis for slow modes. Here, we apply the approach to an analytically solvable model system as well as on three protein systems of different complexities. We consistently demonstrate that the inferred indicator functions can significantly improve the estimation of the leading eigenvalues of Koopman operators and correctly identify key states and transition time scales of stochastic systems, even when good observables are not known a priori.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning-Guided Identification of PET Hydrolases from Natural Diversity

The enzymatic depolymerization of poly(ethylene terephthalate) (PET) is emerging as a leading chemical recycling technology for waste polyester. As part of this endeavor, new candidate enzymes identified from natural diversity can serve as useful starting points for enzyme evolution and engineering. In this study, we improved upon HMM searches by applying an iterative machine learning strategy to identify 400 putative PET-degrading enzymes (PET hydrolases) from naturally occurring homologs. Using high-throughput (HTP) experimental techniques, we successfully expressed and purified >200 enzyme candidates and assayed them for PET hydrolysis activity as a function of pH, temperature, and substrate crystallinity. From this library, we discovered 91 previously unknown PET hydrolases, 35 of which retain activity at pH 4.5 on crystalline material, which are conditions relevant to developing more efficient commercial processes. Notably, four enzymes showed equal to or higher activity than LCC-ICCG, a benchmark PET hydrolase, at this challenging condition in our screening assay, and 11 of which have pH optima <7. Using these data, we identified regions of PETases statistically correlated to activity at lower pH. We additionally investigated the effect of condition-specific activity data on trained machine learning predictors and found a precision (putative hit rate) improvement of up to 30% compared to a Hidden Markov Model alone. Our findings show that by pointing enzyme discovery toward conditions of interest with multiple rounds of experimental and machine learning, we can discover large sets of active enzymes and explore factors associated with activity at those conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Engineering and Application of a Thermostable MHETase for PET Depolymerization

Enzymatic hydrolysis of poly(ethylene terephthalate) (PET) releases mono(2-hydroxyethyl) terephthalate (MHET) as a major product, the accumulation of which can prolong reactor residence times and complicate downstream monomer separations. The use of a MHETase enzyme can enable MHET hydrolysis to the monomers, terephthalic acid and ethylene glycol, but industrial PETases typically operate at thermophilic temperatures and the well-known MHETase from Ideonella sakaiensis is a mesophilic enzyme, thus warranting the development of thermophilic MHETases. Here, we characterize thermostable MHET-active enzymes from a natural diversity screen by applying a hidden Markov model based on the previously reported, archaeal ferulic acid esterase, PET46. We identified enzymes with higher thermostability than PET46 and quantified their MHETase activity in reactions at 70 °C. The crystal structure of MHT077, the homologue with the highest MHETase activity and an apparent melting temperature (T m,app ) of 94.6 °C, informed site saturation mutagenesis in the active site and lid-domain interface. MHT077 exhibited a ∼100-fold slower unfolding rate at 65 °C than PET46, indicating substantially greater kinetic stability. In parallel, we applied evolution-informed design, a probabilistic model that leverages coevolutionary patterns in large multiple sequence alignments, to improve the activity and thermostability of five ferulic acid esterases. One design, EV-MHT043–5 was identified with a comparable thermostability (T m,app = 96.1 °C) and a 3-fold improvement in its MHETase activity relative to the wildtype enzyme, MHT043. Combination variants of beneficial mutations were screened and afforded a variant, MHT077 LFK , which reduced MHET accumulation in bioreactor experiments with postconsumer PET waste. Overall, this study expands the known MHET-hydrolyzing protein scaffolds available for enzymatic PET recycling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Unsupervised Clustering of Microseismic Events and Focal Mechanism Analysis at the CO 2 Injection Site in Decatur, Illinois

Characterization of induced microseismicity at a carbon dioxide (CO 2 ) storage site is critical for preserving reservoir integrity and mitigating seismic hazards. We apply a multilevel machine learning (ML) approach that combines the nonnegative matrix factorization and hidden Markov model to extract spectral representations of microseismic events and cluster them to identify seismic patterns at the Illinois Basin-Decatur Project. Unlike traditional waveform correlation methods, this approach leverages spectral characteristics of first arrivals to improve event classification and detect previously undetected planes of weakness. By integrating ML-based clustering with focal mechanism analysis, we resolve small-scale fault structures that are below the detection limits of conventional seismic imaging. Our findings reveal temporal bursts of microseismicity associated with brittle failure, providing insights into the spatio-temporal evolution of fault reactivation during CO 2 injection. This approach enhances seismic monitoring capabilities at CO 2 injection sites by improving fault characterization beyond the resolution of standard geophysical surveys.

Willis, Rachel Marie [Sandia National Laboratories↗

Targeted assemblies of cas1 suggest CRISPR-Cas’s response to soil warming

Abstract There is an increasing interest in the clustered regularly interspaced short palindromic repeats CRISPR-associated protein (CRISPR-Cas) system to reveal potential virus–host dynamics. The universal and most conserved Cas protein, cas1 is an ideal marker to elucidate CRISPR-Cas ecology. We constructed eight Hidden Markov Models (HMMs) and assembled cas1 directly from metagenomes by a targeted-gene assembler, Xander, to improve detection capacity and resolve the diverse CRISPR-Cas systems. The eight HMMs were first validated by recovering all 17 cas1 subtypes from the simulated metagenome generated from 91 prokaryotic genomes across 11 phyla. We challenged the targeted method with 48 metagenomes from a tallgrass prairie in Central Oklahoma recovering 3394 cas1. Among those, 88 were near full length, 5 times more than in de-novo assemblies from the Oklahoma metagenomes. To validate the host assignment by cas1, the targeted-assembled cas1 was mapped to the de-novo assembled contigs. All the phylum assignments of those mapped contigs were assigned independent of CRISPR-Cas genes on the same contigs and consistent with the host taxonomies predicted by the mapped cas1. We then investigated whether 8 years of soil warming altered cas1 prevalence within the communities. A shift in microbial abundances was observed during the year with the biggest temperature differential (mean 4.16 °C above ambient). cas1 prevalence increased and even in the phyla with decreased microbial abundances over the next 3 years, suggesting increasing virus–host interactions in response to soil warming. This targeted method provides an alternative means to effectively mine cas1 from metagenomes and uncover the host communities.

54 ENVIRONMENTAL SCIENCES↗

DIPS-Plus: The enhanced database of interacting protein structures for interface prediction

Abstract In this work, we expand on a dataset recently introduced for protein interface prediction (PIP), the Database of Interacting Protein Structures (DIPS), to present DIPS-Plus, an enhanced, feature-rich dataset of 42,112 complexes for machine learning of protein interfaces. While the original DIPS dataset contains only the Cartesian coordinates for atoms contained in the protein complex along with their types, DIPS-Plus contains multiple residue-level features including surface proximities, half-sphere amino acid compositions, and new profile hidden Markov model (HMM)-based sequence features for each amino acid, providing researchers a curated feature bank for training protein interface prediction methods. We demonstrate through rigorous benchmarks that training an existing state-of-the-art (SOTA) model for PIP on DIPS-Plus yields new SOTA results, surpassing the performance of some of the latest models trained on residue-level and atom-level encodings of protein complexes to date.

59 BASIC BIOLOGICAL SCIENCES↗

Behavioral state resource selection in invasive wild pigs in the Southeastern United States

Elucidating correlations between wild pig ( Sus scrofa ) behavior and landscape attributes can aid in the advancement of management strategies for controlling populations. Using GPS data from 49 wild pigs in the southeastern U.S., we used hidden Markov models to define movement path characteristics and assign behaviors (e.g., resting, foraging, travelling). We then explored the connection between these behaviors and resource selection for both sexes between two distinct seasons based on forage availability (i.e., low forage, high forage). Females demonstrated a crepuscular activity pattern in the high-forage season and a variable pattern in the low-forage season, while males exhibited nocturnal activity patterns across both seasons. Wild pigs selected for bottomland hardwoods and dense canopy cover in all behavioral states in both seasons. Males selected for diversity in vegetation types while foraging in the low-forage season compared to the high-forage season and demonstrated an increased use of linear anthropogenic features across seasons while traveling. Wild pigs can establish populations and home ranges in an array of landscapes, but our results demonstrate male and female pigs exhibit clear differences in movement behavior and there are key resources associated with common behaviors that can be targeted to improve the efficiency of management programs.

54 ENVIRONMENTAL SCIENCES↗

Sequence-structure-function characterization of the emerging tetracycline destructase family of antibiotic resistance enzymes

Tetracycline destructases (TDases) are flavin monooxygenases which can confer resistance to all generations of tetracycline antibiotics. The recent increase in the number and diversity of reported TDase sequences enables a deep investigation of the TDase sequence-structure-function landscape. Here, we evaluate the sequence determinants of TDase function through two complementary approaches: (1) constructing profile hidden Markov models to predict new TDases, and (2) using multiple sequence alignments to identify conserved positions important to protein function. Using the HMM-based approach we screened 50 high-scoring candidate sequences in Escherichia coli, leading to the discovery of 13 new TDases. The X-ray crystal structures of two new enzymes from Legionella species were determined, and the ability of anhydrotetracycline to inhibit their tetracycline-inactivating activity was confirmed. Using the MSA-based approach we identified 31 amino acid positions 100% conserved across all known TDase sequences. The roles of these positions were analyzed by alanine-scanning mutagenesis in two TDases, to study the impact on cell and in vitro activity, structure, and stability. These results expand the diversity of TDase sequences and provide valuable insights into the roles of important residues in TDases, and flavin monooxygenases more broadly.

60 APPLIED LIFE SCIENCES↗

Uncovering heterogeneous intercommunity disease transmission from neutral allele frequency time series

The COVID-19 pandemic has underscored the need for accurate epidemic forecasting to predict pathogen spread, evolution, and evaluate intervention strategies. Forecast reliability hinges on detailed knowledge of disease transmission across population segments, which may be inferred from contact surveys or mobility data. However, these indirect approaches make it difficult to estimate rare transmissions between socially or geographically distant communities. We show that the steep ramp-up of genome sequencing surveillance during the pandemic can be leveraged to directly identify transmission patterns between geographically defined communities. Our approach uses a hidden Markov model to infer the fraction of infections a community imports from others based on how rapidly allele frequencies in the focal community converge to those in the donor communities. Applying this method to SARS-CoV-2 sequencing data from England and the United States, we uncover networks of intercommunity transmission that reflect geographical relationships while exposing significant long-range interactions. The scaling of importation rate with distance is consistent across both countries, yet weaker than expected based on mobility data, highlighting limitations of indirect inference. We show that transmission patterns can change between waves of variants of concern and analyze how the inferred heterogeneity in intercommunity transmission impacts evolutionary forecasts. While applied here to geographically defined communities, our approach could be applied to those defined by other traits (e.g., age, socioeconomic status), provided time-series data can be stratified accordingly. Overall, our study highlights population genomic time series data as a crucial record of epidemiological interactions, which can be deciphered using tree-free inference methods.

Okada, Takashi [Department of Physics; University ↗

Sensitive and error-tolerant annotation of protein-coding DNA with BATH

We present BATH, a tool for highly sensitive annotation of protein-coding DNA based on direct alignment of that DNA to a database of protein sequences or profile hidden Markov models (pHMMs). BATH is built on top of the HMMER3 code base, and simplifies the annotation workflow for pHMM-based translated sequence annotation by providing a straightforward input interface and easy-to-interpret output. BATH also introduces novel frameshift-aware algorithms to detect frameshift-inducing nucleotide insertions and deletions (indels). BATH matches the accuracy of HMMER3 for annotation of sequences containing no errors, and produces superior accuracy to all tested tools for annotation of sequences containing nucleotide indels. These results suggest that BATH should be used when high annotation sensitivity is required, particularly when frameshift errors are expected to interrupt protein-coding regions, as is true with long-read sequencing data and in the context of pseudogenes.

59 BASIC BIOLOGICAL SCIENCES↗

efam: an e xpanded, metaproteome-supported HMM profile database of viral protein fam ilies

Viruses infect, reprogram and kill microbes, leading to profound ecosystem consequences, from elemental cycling in oceans and soils to microbiome-modulated diseases in plants and animals. Although metagenomic datasets are increasingly available, identifying viruses in them is challenging due to poor representation and annotation of viral sequences in databases. Here, we establish efam, an expanded collection of Hidden Markov Model (HMM) profiles that represent viral protein families conservatively identified from the Global Ocean Virome 2.0 dataset. This resulted in 240 311 HMM profiles, each with at least 2 protein sequences, making efam >7-fold larger than the next largest, pan-ecosystem viral HMM profile database. Adjusting the criteria for viral contig confidence from ‘conservative’ to ‘eXtremely Conservative’ resulted in 37 841 HMM profiles in our efam-XC database. To assess the value of this resource, we integrated efam-XC into VirSorter viral discovery software to discover viruses from less-studied, ecologically distinct oxygen minimum zone (OMZ) marine habitats. This expanded database led to an increase in viruses recovered from every tested OMZ virome by ~24% on average (up to ~42%) and especially improved the recovery of often-missed shorter contigs (<5 kb). Additionally, to help elucidate lesser-known viral protein functions, we annotated the profiles using multiple databases from the DRAM pipeline and virion-associated metaproteomic data, which doubled the number of annotations obtainable by standard, single-database annotation approaches. Together, these marine resources (efam and efam-XC) are provided as searchable, compressed HMM databases that will be updated bi-annually to help maximize viral sequence discovery and study from any ecosystem.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NMPFamsDB: a database of novel protein families from microbial metagenomes and metatranscriptomes

Abstract The Novel Metagenome Protein Families Database (NMPFamsDB) is a database of metagenome- and metatranscriptome-derived protein families, whose members have no hits to proteins of reference genomes or Pfam domains. Each protein family is accompanied by multiple sequence alignments, Hidden Markov Models, taxonomic information, ecosystem and geolocation metadata, sequence and structure predictions, as well as 3D structure models predicted with AlphaFold2. In its current version, NMPFamsDB hosts over 100 000 protein families, each with at least 100 members. The reported protein families significantly expand (more than double) the number of known protein sequence clusters from reference genomes and reveal new insights into their habitat distribution, origins, functions and taxonomy. We expect NMPFamsDB to be a valuable resource for microbial proteome-wide analyses and for further discovery and characterization of novel functions. NMPFamsDB is publicly available in http://www.nmpfamsdb.org/ or https://bib.fleming.gr/NMPFamsDB.

59 BASIC BIOLOGICAL SCIENCES↗

metagRoot: a comprehensive database of protein families associated with plant root microbiomes

The plant root microbiome is vital in plant health, nutrient uptake, and environmental resilience. To explore and harness this diversity, we present metagRoot, a specialized and enriched database focused on the protein families of the plant root microbiome. MetagRoot integrates metagenomic, metatranscriptomic, and reference genome-derived protein data to characterize 71 091 enriched protein families, each containing at least 100 sequences. These families are annotated with multiple sequence alignments, CRISPR elements, hidden Markov models, taxonomic and functional classifications, ecosystem and geolocation metadata, and predicted 3D structures using AlphaFold2. MetagRoot is a powerful tool for decoding the molecular landscape of root-associated microbial communities and advancing microbiome-informed agricultural practices by enriching protein family information with ecological and structural context. The database is available at https://pavlopoulos-lab.org/metagroot/ or https://www.metagroot.org.

Chasapi, Maria N↗

CasCollect: targeted assembly of CRISPR-associated operons from high-throughput sequencing data

Abstract CRISPR arrays and CRISPR-associated (Cas) proteins comprise a widespread adaptive immune system in bacteria and archaea. These systems function as a defense against exogenous parasitic mobile genetic elements that include bacteriophages, plasmids and foreign nucleic acids. With the continuous spread of antibiotic resistance, knowledge of pathogen susceptibility to bacteriophage therapy is becoming more critical. Additionally, gene-editing applications would benefit from the discovery of new cas genes with favorable properties. While next-generation sequencing has produced staggering quantities of data, transitioning from raw sequencing reads to the identification of CRISPR/Cas systems has remained challenging. This is especially true for metagenomic data, which has the highest potential for identifying novel cas genes. We report a comprehensive computational pipeline, CasCollect, for the targeted assembly and annotation of cas genes and CRISPR arrays—even isolated arrays—from raw sequencing reads. Benchmarking our targeted assembly pipeline demonstrates significantly improved timing by almost two orders of magnitude compared with conventional assembly and annotation, while retaining the ability to detect CRISPR arrays and cas genes. CasCollect is a highly versatile pipeline and can be used for targeted assembly of any specialty gene set, reconfigurable for user provided Hidden Markov Models and/or reference nucleotide sequences.

Podlevsky, Joshua D.↗

Labeling sequential data from noisy annotations

Crowdsourcing algorithms often work under the assumption that the data samples are independent. Recent work has shown that data dependence, such as temporal correlations in sequential data, can be leveraged to improve the label quality. Existing methods that exploit this special structure rely on third-order statistics of the annotator outputs to ensure the identifiability of key latent parameters, which are costly to acquire. This work proposes an approach for integrating crowdsourced annotations under the Dawid-Skene/Hidden Markov Model (DS-HMM) for sequential data based on second-order statistics, which naturally enjoys a lower sample complexity. An effective algorithm is proposed to tackle the challenging optimization problem associated with the proposed estimator. Numerical experiments showcase the effectiveness of the data labeling paradigm.

Marrinan, Timothy P.↗

On the Performance of Detecting Injection of Fabricated Messages into the CAN Bus

There have been several public demonstrations of attacks on connected vehicles showing the ability of an attacker to take control of a targeted vehicle by injecting messages into their Controller Area Network (CAN) bus. In this article, using injected speed reading and Revolutions Per Minute (RPM) reading messages in in-motion vehicle, we examine the ability of the Pearson correlation and the unsupervised learning methods k-means clustering and Hidden Markov Model (HMM) to differentiate ’no-attack’ and ’under-attack’ states of the given vehicle. We found that the Pearson correlation distinguishes the two states, the k-means clustering method has an acceptable accuracy but high false positive rate and HMM detects attacks with acceptable detection rate but has a high false positive in detecting attacks from speed readings when there is no attack. The accuracy of these unsupervised learning methods are comparable to the ones of the supervised learning methods used by CAN bus Intrusion Detection System (IDS) suppliers. In addition, the article shows that studying CAN anomaly detection techniques using off-vehicle test facilities may not properly evaluate the performance of the detection techniques. Finally, the results suggest using other features besides the data content of the CAN messages and integrate knowledge about how the Electronic Control Units (ECUs) collaborate in building effective techniques for the detection of injection of fabricated message attacks.

97 MATHEMATICS AND COMPUTING↗

A Stochastic Multi-Criteria Decision-Making Algorithm for Dynamic Load Prioritization in Grid-Interactive Efficient Buildings

Increasing deployment of advanced sensing, controls, and communication infrastructure enables buildings to provide services to the power grid, leading to the concept of grid-interactive efficient buildings. Since occupant activities and preferences primarily drive the availability and operational flexibility of building devices, there is a critical need to develop occupant-centric approaches that prioritize devices for providing grid services, while maintaining the desired end-use quality of service. In this paper, we present a decision-making framework that facilitates a building owner/operator to effectively prioritize loads for curtailment service under uncertainties, while minimizing any adverse impact on the occupants. The proposed framework uses a stochastic (Markov) model to represent the probabilistic behavior of device usage from power consumption data, and a load prioritization algorithm that dynamically ranks building loads using a stochastic multi-criteria decision-making algorithm. The proposed load prioritization framework is illustrated via numerical simulations in a residential building use-case, including plug-loads, air-conditioners, and plug-in electric vehicle chargers, in the context of load curtailment as a grid service. Suitable metrics are proposed to evaluate the closed-loop performance of the proposed prioritization algorithm under various scenarios and design choices. Scalability of the proposed algorithm is established via computational analysis, while time-series plots are used for intuitive explanation of the ranking choices.

24 POWER TRANSMISSION AND DISTRIBUTION↗