Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine-learned interatomic potentials by active learning: amorphous and liquid hafnium dioxide

We propose an active learning scheme for automatically sampling a minimum number of uncorrelated configurations for fitting the Gaussian Approximation Potential (GAP). Our active learning scheme consists of an unsupervised machine learning (ML) scheme coupled with a Bayesian optimization technique that evaluates the GAP model. We apply this scheme to a Hafnium dioxide (HfO 2 ) dataset generated from a "melt-quench" ab initio molecular dynamics (AIMD) protocol. Our results show that the active learning scheme, with no prior knowledge of the dataset, is able to extract a configuration that reaches the required energy fit tolerance. Further, molecular dynamics (MD) simulations performed using this active learned GAP model on 6144 atom systems of amorphous and liquid state elucidate the structural properties of HfO 2 with near ab initio precision and quench rates (i.e., 1.0 K/ps) not accessible via AIMD. The melt and amorphous X-ray structural factors generated from our simulation are in good agreement with experiment. In addition, the calculated diffusion constants are in good agreement with previous ab initio studies.

36 MATERIALS SCIENCE↗

Unsupervised learning for identifying events in active target experiments

This article presents novel applications of unsupervised machine learning methods to the problem of event separation in an active target detector, the Active-Target Time Projection Chamber (AT-TPC). The overarching goal is to group similar events in the early stages of the data analysis, thereby improving efficiency by limiting the computationally expensive processing of unnecessary events. The application of unsupervised clustering algorithms to the analysis of two-dimensional projections of particle tracks from a resonant proton scattering experiment on 46 Ar is introduced. We explore the performance of autoencoder neural networks and a pre-trained VGG16 Simonyan and Zisserman (2015) convolutional neural network. We study clustering performance on both data from a simulated 46 Ar experiment, and real events from the AT-TPC detector. We find that a -means algorithm applied to simulated data in the VGG16 latent space forms almost perfect clusters. Additionally, the VGG16+-means approach finds high purity clusters of proton events for real experimental data. Here, we also explore the application of clustering the latent space of autoencoder neural networks for event separation. While these networks show strong performance, they suffer from high variability in their results.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine Learning-Enabled Quantitative Analysis of Optically Obscure Scratches on Nickel-Plated Additively Manufactured (AM) Samples

Additively manufactured metal components often have rough and uneven surfaces, necessitating post-processing and surface polishing. Hardness is a critical characteristic that affects overall component properties, including wear. This study employed K-means unsupervised machine learning to explore the relationship between the relative surface hardness and scratch width of electroless nickel plating on additively manufactured composite components. The Taguchi design of experiment (TDOE) L9 orthogonal array facilitated experimentation with various factors and levels. Initially, a digital light microscope was used for 3D surface mapping and scratch width quantification. However, the microscope struggled with the reflections from the shiny Ni-plating and scatter from small scratches. To overcome this, a scanning electron microscope (SEM) generated grayscale images and 3D height maps of the scratched Ni-plating, thus enabling the precise characterization of scratch widths. Optical identification of the scratch regions and quantification were accomplished using Python code with a K-means machine-learning clustering algorithm. The TDOE yielded distinct Ni-plating hardness levels for the nine samples, while an increased scratch force showed a non-linear impact on scratch widths. The enhanced surface quality resulting from Ni coatings will have significant implications in various industrial applications, and it will play a pivotal role in future metal and alloy surface engineering.

36 MATERIALS SCIENCE↗

Latent Representation Learning for Structural Characterization of Catalysts

Supervised machine learning-enabled mapping of the X-ray absorption near edge structure (XANES) spectra to local structural descriptors offers new methods for understanding the structure and function of working nanocatalysts. We briefly summarize a status of XANES analysis approaches by supervised machine learning methods. We present an example of an autoencoder-based, unsupervised machine learning approach for latent representation learning of XANES spectra. This new approach produces a lower-dimensional latent representation, which retains a spectrum–structure relationship that can be eventually mapped to physicochemical properties. Furthermore, the latent space of the autoencoder also provides a pathway to interpret the information content “hidden” in the X-ray absorption coefficient. Our approach (that we named latent space analysis of spectra, or LSAS) is demonstrated for the supported Pd nanoparticle catalyst studied during the formation of Pd hydride. By employing the low-dimensional representation of Pd K-edge XANES, the LSAS method was able to isolate the key factors responsible for the observed spectral changes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Creating ground truth for nanocrystal morphology: a fully automated pipeline for unbiased transmission electron microscopy analysis

Control over colloidal nanocrystal morphology (size, size distribution, and shape) is important for tailoring the functionality of individual nanocrystals and their ensemble behavior. Despite this, traditional methods to quantify nanocrystal morphology are laborious. New developments in automated morphology classification will accelerate these analyses but the assessment of machine learning models is limited by human accuracy for ground truth, causing even unsupervised machine learning models to have inherent bias. Herein, we introduce synthetic image rendering to solve the ground truth problem of nanocrystal morphology classification. By simulating 2D images of nanocrystal shapes via a function of high-dimensional parameter space, we trained a convolutional neural network to link unique morphologies to their simulated parameters, defining nanocrystal morphology quantitatively rather than qualitatively. An automated pipeline then processes, quantitatively defines, and classifies nanocrystal morphology from experimental transmission electron microscopy (TEM) images. Using improved computer vision techniques, 42,650 nanocrystals were identified, assessed, and labeled with quantitative parameters, offering a 600-fold improvement in efficiency over best-practice manual measurements. Further, a classification algorithm was trained with a prediction accuracy of 99.5%, which can successfully analyze a range of concave, convex, and irregular nanocrystal shapes. The resulting pipeline was applied to differentiating two syntheses of nominally cuboidal CsPbBr 3 nanocrystals and uniquely classifying binary nickel sulfide nanocrystal phase based on morphology. This pipeline provides a simple, efficient, and unbiased method to quantify nanocrystal morphology and represents a practical route to construct large datasets with an absolute ground truth for training unbiased morphology-based machine learning algorithms.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Predicting the Seawater Chemistry of an Ocean World Using Machine Learning on Isotopic Measurements of Volatile CO2

Introduction: Given the long time intervals required for data transmission to and from ocean worlds targets, low bandwidth for data transmission, time required for data processing and analysis, and potentially extreme radiation environments (e.g., Europa), it is clear that ocean worlds missions will need more autonomous flight instruments and software in order to achieve established science goals. Protracted time intervals for data analysis (e.g., Europa Lander) strongly motivates the development of rapid, consistent and streamlined methods for interpreting data from flight mass spectrometers to e.g., determine how mass spectra from a plume or surface liquid/ice relates to the surface/subsurface. Since mass spectrometry also has the potential to correctly identify biosignatures[1], it is imperative that such methods for interpreting data are consistent and accurate. We used 848 isotope ratio mass spectra from laboratory analyses of CO2 that interacted with ocean worlds-relevant seawaters as a ‘training’ dataset for ‘unsupervised’ machine learning. In unsupervised learning, characteristics of the data are not labeled or linked, and any similarities found only result from the neural network. CO2 isotopologues analyzed for this dataset mimic the remote measurements of CO2 by a flight mass spectrometer, and are detailed in Theiling [2]. From this dataset, we used measured features of the spectra, such as retention time, intensity, and (isotopologue) mass ratios as inputs for our autoencoder neural network. Our neural network was trained to find similarities in these and other spectral features for seawaters of a particular composition and amount of initial CO2. Successful training then created an output of these similarities for various seawaters, which included MgSO4, Na2SO4, NaCl, MgCl2, KCl, and NaHCO3, and combinations of these salts. We then applied dimensionality reduction techniques such as Principal Component Analysis (PCA), T-Distributed Stochastic Neighbor Embedding (TSNE), and Uniform Manifold Approximation and Projection (UMAP) to demonstrate latent data features as a two-dimensional projection in a unitless, high-dimensional space. In this projection, a data point represents the combined effect of spectral features such as intensity, retention time, and isotope ratio. Our initial UMAP demonstrates data clustering (organization of the data by the neural network) based on the amount of CO2 that had initially interacted with each seawater. Further training using more ‘supervised’ learning techniques demonstrate strong clustering of preliminary data based on initial CO2 concentration, seawater chemical composition, and ionic strength (salinity). Our preliminary work therefore suggests that machine learning has the potential to identify compositional variants of an ocean world seawater based on mass spectra from volatile CO2 measurements. Acknowledgments: This work was funded through a Strategic Task Group at NASA Goddard Space Flight Center. The training dataset was collected through funding from the Oklahoma Space Grant Consortium. References: [1] Pappalardo, R. et al. (2013) Astrobiology, 13, 740–773. [2] Theiling (2020) Icarus, 114216.

Europa↗

Classical versus quantum models in machine learning: insights from a finance application

Although several models have been proposed towards assisting machine learning (ML) tasks with quantum computers, a direct comparison of the expressive power and efficiency of classical versus quantum models for datasets originating from real-world applications is one of the key milestones towards a quantum ready era. Here, we take a first step towards addressing this challenge by performing a comparison of the widely used classical ML models known as restricted Boltzmann machines (RBMs), against a recently proposed quantum model, now known as quantum circuit Born machines (QCBMs). Both models address the same hard tasks in unsupervised generative modeling, with QCBMs exploiting the probabilistic nature of quantum mechanics and a candidate for near-term quantum computers, as experimentally demonstrated in three different quantum hardware architectures to date. To address the question of the performance of the quantum model on real-world classical data sets, we construct scenarios from a probabilistic version out of the well-known portfolio optimization problem in finance, by using time-series pricing data from asset subsets of the S&P500 stock market index. It is remarkable to find that, under the same number of resources in terms of parameters for both classical and quantum models, the quantum models seem to have superior performance on typical instances when compared with the canonical training of the RBMs. Our simulations are grounded on a hardware efficient realization of the QCBMs on ion-trap quantum computers, by using their native gate sets, and therefore readily implementable in near-term quantum devices.

97 MATHEMATICS AND COMPUTING↗

Selecting representative geological realizations to model subsurface CO 2 storage under uncertainty

Carbon capture and storage (CCS) is one of the quickest and most effective solutions for reducing carbon emissions. The majority of subsurface storage occurs in saline aquifers, for which geological information is lacking which in turn results in geological uncertainty. To evaluate uncertainty in CO 2 injection projections, the use of multiple geological realizations (GRs) has been practiced very commonly. In this approach, hundreds or thousands of high-resolution GRs is used that quickly becomes computationally expensive. This issue can be addressed with representative geological realizations (RGRs) that preserve the uncertainty domain of the ensemble GRs. Here, in this study, we propose the use of unsupervised machine learning (UML) frameworks, including dissimilarity measurement, dimensionality reduction, clustering and sampling algorithms ta select a predetermined number of RGRs. We compare the simulation outputs of the RGR sets and the ensemble using the Kolmogorov–Smirnov (KS) test to select the best UML. The UML frameworks and their associated selection processes are evaluated using a saline aquifer with a single CO 2 injection well and 200 GRs with varying uncertain petrophysical characteristics. The best UML framework is selected to use only 5% of the GRs while maintaining the uncertainty domain of the ensemble GRs. In addition, the best UML framework is tested using a saline aquifer with three CO 2 injection wells and varied GRs. The results show that our proposed UML framework can be used to choose RGRs, capturing the whole uncertainty domain. Our approach leads to a significant reduction in the computational cost associated with scenario testing, decision-making, and development planning for CO 2 storage sites under geological uncertainty.

58 GEOSCIENCES↗

FIB-ToF-SIMS characterization of irradiated U-10Zr

Post-irradiation examination (PIE) is critical for the performance assessment and qualification of nuclear fuels. Secondary ion mass spectrometry (SIMS) is a powerful materials characterization technique that allows for elemental and isotopic mapping with a depth resolution greater than EDS and EPMA. However, it has not yet been applied to PIE of metallic nuclear fuel. Here, in this work, we characterize an fast neutron spectrum irradiated U-10Zr fuel sample using a time-of-flight SIMS (ToF-SIMS) system connected to a FIB/SEM system, which allows for flexible sample analysis compared to a dedicated ToF-SIMS instrument. Analysis of the resulting hyperspectral micrograph data was aided by the development of an unsupervised machine learning (ML) algorithm that iterates on existing methods to segment the 3D micrographic datasets based on the similarity of mass spectra. The results showed that the FIB-ToF-SIMS instrument was potentially capable of spatially resolving closed fission gas bubbles in 3D by continued ion sputtering of the analyzed volume. Additionally, the ML algorithm proved useful in revealing the chemical segregation of light fission products (those with an atomic mass between approximately 85–105 amu, such as ruthenium and rhodium) plus matrix zirconium, heavy fission products (those with an atomic mass between approximately 135–150 amu, such as the lanthanides) and uranium. Future studies are planned to conduct FIB-ToF-SIMS analysis on more irradiated U-Zr samples to study the constituent redistribution.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Remote Sensing-Informed Zonation for Understanding Snow, Plant and Soil Moisture Dynamics within a Mountain Ecosystem

In the headwater catchments of the Rocky Mountains, plant productivity and its dynamics are largely dependent upon water availability, which is influenced by changing snowmelt dynamics associated with climate change. Understanding and quantifying the interactions between snow, plants and soil moisture is challenging, since these interactions are highly heterogeneous in mountainous terrain, particularly as they are influenced by microtopography within a hillslope. Recent advances in satellite remote sensing have created an opportunity for monitoring snow and plant dynamics at high spatiotemporal resolutions that can capture microtopographic effects. In this study, we investigate the relationships among topography, snowmelt, soil moisture and plant dynamics in the East River watershed, Crested Butte, Colorado, based on a time series of 3-meter resolution PlanetScope normalized difference vegetation index (NDVI) images. To make use of a large volume of high-resolution time-lapse images (17 images total), we use unsupervised machine learning methods to reduce the dimensionality of the time lapse images by identifying spatial zones that have characteristic NDVI time series. We hypothesize that each zone represents a set of similar snowmelt and plant dynamics that differ from other identified zones and that these zones are associated with key topographic features, plant species and soil moisture. We compare different distance measures (Ward and complete linkage) to understand the effects of their influence on the zonation map. Results show that the identified zones are associated with particular microtopographic features; highly productive zones are associated with low slopes and high topographic wetness index, in contrast with zones of low productivity, which are associated with high slopes and low topographic wetness index. The zones also correspond to particular plant species distributions; higher forb coverage is associated with zones characterized by higher peak productivity combined with rapid senescence in low moisture conditions, while higher sagebrush coverage is associated with low productivity and similar senescence patterns between high and low moisture conditions. In addition, soil moisture probe and sensor data confirm that each zone has a unique soil moisture distribution. This cluster-based analysis can tractably analyze high-resolution time-lapse images to examine plant-soil-snow interactions, guide sampling and sensor placements and identify areas likely vulnerable to ecological change in the future.

54 ENVIRONMENTAL SCIENCES↗

Uranium Oxide Synthetic Pathway Discernment through Unsupervised Morphological Analysis

We present a novel unsupervised machine learning method for quantitative representation of scanning electron micrographs and its applications and performance for nuclear forensic analysis of uranium ore concentrates. The method uses a vector quantizing variational autoencoder followed by a histogram operation to encode a micrograph into a single dimensional representation, called the latent vector. The method requires no extant labeling of the data and can be applied over large datasets of micrographs with minimal human interaction. The representations generated are broadly descriptive of each micrograph and the microstructure of the material imaged. In the case of uranium ore concentrate analysis, the representations were amenable to processing reagent and ore concentrate species classification with accuracy of 81:8%, which is competitive with state-of-the-art supervised networks. The representations were also used to classify previously unseen processing routes, were able to classify imaging parameters such as magnification (to 76:0% accuracy), were able to classify fine grained process parameters such as calcining temperature (to 74:4% accuracy), and their informatic properties indicate that they are generally descriptive of the image represented. This method can be applied across microstructure analysis fields to perform quantitative analysis without the need for labor intensive and possibly biased human analysis.

Scanning Electron Microscopy, Vector Quantizing Va↗

Machine learning to identify geologic factors associated with production in geothermal fields: A casestudy using 3D geologic data, Brady geothermal field, Nevada

In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in the Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fluid flow pathways are relatively rare in the subsurface but are critical components of hydrothermal systems like Brady and many other types of fluid flow systems in fractured rock. The ML method, non-negative matrix factorization with k-means clustering (NMFk), is applied to a library of fourteen 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate the macro-scale faults and a local step-over in the fault system preferentially occur along with production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate 1) the specific geologic controls on the Brady hydrothermal system and 2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes.

15 GEOTHERMAL ENERGY↗

Machine Learning to Identify Geologic Factors Associated with Production in Geothermal Fields: A Case-Study Using 3D Geologic Data from Brady Geothermal Field and NMFk

In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fuid-fow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fuid-fow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k-means clustering (NMFk), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes. This submission includes the published journal article detailing this work, the published 3D geologic map of the Brady Geothermal Area used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity, 3D well data, along which geologic data were sampled for PCA analyses, and associated metadata file. This work was done using the GeoThermalCloud framework, which is part of SmartTensors (both are linked below).

15 GEOTHERMAL ENERGY↗

Machine Learning Reveals Memory of the Parent Phases in Ferroelectric Relaxors Ba(Ti$_{1-x}$,Zr x )O 3

Machine learning has been establishing its potential in multiple areas of condensed matter physics and materials science. Here, in this work, an unsupervised machine learning workflow is developed and used within a framework of first-principles-based atomistic simulations to investigate phases, phase transitions, and their structural origins in ferroelectric relaxors, Ba(Ti 1-x ,Zr x )O 3 . The applicability of the workflow is first demonstrated to identify phases and phase transitions in the parent compound, a prototypical ferroelectric BaTiO 3 . Then the workflow is applied for Ba(Ti 1-x ,Zrx)O 3 with x ≤ 0.25 to reveal i) that some of the compounds bear a subtle memory of BaTiO 3 phases beyond the point of the pinched phase transition, which could contribute to their enhanced electromechanical response; ii) the existence of peculiar phases with delocalized precursors of nanodomains—likely candidates for the controversial polar nanoregions; and iii) nanodomain phases for the largest concentrations of x.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Machine-learning predictions of the shale wells’ performance

The ultra-low permeability nature of shale reservoirs leads to an extended linear flow and necessitates horizontal wells with multi-stage engineered fractures to efficiently extract hydrocarbons resources. These artificially-generated and naturally-occurring fractures form complex networks that create complex flow regimes which control oil production. These fractures are neither identical nor equally-spaced, which leads to a production profile with a masked onset of the boundary-dominated flow. The combination of the extended linear flow with the indeterminate onset of the boundary-dominated flow challenges the current deterministic analytic approaches to forecast the estimated ultimate recovery (EUR). In this work, we propose a novel machine-learning approach which overcomes these challenges and provides reliable EUR estimates based on field-wide analyses. We implement a novel unsupervised machine learning (ML) methodology, which allows for automatic identification of the optimal number of features (signals) present in the data based on non-negative matrix/tensor factorization coupled with k-means clustering incorporating regularization and physics constraints. In the presented analyses, the input data to the ML algorithm is the available (public) production history from the field collected at existing unconventional reservoirs. We validate our approach through hindcasting of the production data, where we achieved an excellent agreement. In addition, our approach is able to identify the poorly-performing wells, which could benefit from early refracing. Our approach provides fast and accurate estimations of the well performance without presumptions about the state of the well or the flow regime.

03 NATURAL GAS↗

Machine learning to identify geologic factors associated with production in geothermal fields: a case-study using 3D geologic data, Brady geothermal field, Nevada

Abstract In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fluid-flow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fluid-flow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k -means clustering (NMF k ), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes.

58 GEOSCIENCES↗

Hydrogen Bonding Inside Anionic Polymeric Brush Layer: Machine Learning-Driven Exploration of the Relative Roles of the Polymer Steric Effect, Charging, and Type of Screening Counterions

This paper employs a combination of all-atom molecular dynamics (MD) simulations and unsupervised machine learning (ML) for studying the water-water hydrogen bonds (HBs) inside the anionic poly-acrylic acid (PAA) brushes modeled using all-atom MD simulations. PAA brush layer with different charge fraction (f), namely f=0, f=0.25, and f=1, is considered. Water-water interactions, both inside and outside the brush layer, are represented through distinct clusters of tupules of variables representing distances associated with the interacting water molecules. While clusters representing the HBs are present for water inside and outside the brushes, several clusters representing the long-range water-water interactions are missing for the water molecules inside the highly charged (f=1) PAA brushes. More importantly, inside highly charged brushes, the edge of the clusters representing the water-water HBs is progressively shortened, as compared to that in the bulk. Both these results stem from the presence of the PAA brushes imparting the steric effect and the charge effect, or the effect associated with enhanced interactions of water molecules with PE charges and counterions, thereby disrupting the water connectivity. This water-charged-species interaction also increases the water-water HB angle, i.e., makes the water-water HBs less stable inside the highly charged PAA brush layer. The narrowing of the clusters representing the HBs and the alteration of the angle characterizing the HBs confirm that the conditions defining the water-water HBs change inside the PAA brush layer as a function of the charges on the PAA brush layer. Furthermore, we show that the use of the generic definition of HBs, as compared to using our simulation-motivated modified definition of water-water HBs, overpredict the number of water-water HBs inside the PAA brush layer. Finally, we employ this all-atom-MD-ML framework to quantify the effect of other types of screening counterions (Li + , Ca 2+ , and Y 3+ ions) in determining the water-water interactions and water-water HB properties inside the PAA brush layer. Furthermore, the findings of the present study, confirming the weakening of water-water HBs inside the PAA brush layer, points to the possibility that the water molecules will be more available for hydrating the brush layer and counterions, thereby leading to a more pronounced wetting of the PAA brush layer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning uncovers independently regulated modules in the Bacillus subtilis transcriptome

The transcriptional regulatory network (TRN) of Bacillus subtilis coordinates cellular functions of fundamental interest, including metabolism, biofilm formation, and sporulation. Here, we use unsupervised machine learning to modularize the transcriptome and quantitatively describe regulatory activity under diverse conditions, creating an unbiased summary of gene expression. We obtain 83 independently modulated gene sets that explain most of the variance in expression and demonstrate that 76% of them represent the effects of known regulators. The TRN structure and its condition-dependent activity uncover putative or recently discovered roles for at least five regulons, such as a relationship between histidine utilization and quorum sensing. The TRN also facilitates quantification of population-level sporulation states. As this TRN covers the majority of the transcriptome and concisely characterizes the global expression state, it could inform research on nearly every aspect of transcriptional regulation in B. subtilis.

59 BASIC BIOLOGICAL SCIENCES↗