Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “unsupervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Machine Learning Reveals Memory of the Parent Phases in Ferroelectric Relaxors Ba(Ti$_{1-x}$,Zr x )O 3

Machine learning has been establishing its potential in multiple areas of condensed matter physics and materials science. Here, in this work, an unsupervised machine learning workflow is developed and used within a framework of first-principles-based atomistic simulations to investigate phases, phase transitions, and their structural origins in ferroelectric relaxors, Ba(Ti 1-x ,Zr x )O 3 . The applicability of the workflow is first demonstrated to identify phases and phase transitions in the parent compound, a prototypical ferroelectric BaTiO 3 . Then the workflow is applied for Ba(Ti 1-x ,Zrx)O 3 with x ≤ 0.25 to reveal i) that some of the compounds bear a subtle memory of BaTiO 3 phases beyond the point of the pinched phase transition, which could contribute to their enhanced electromechanical response; ii) the existence of peculiar phases with delocalized precursors of nanodomains—likely candidates for the controversial polar nanoregions; and iii) nanodomain phases for the largest concentrations of x.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Hydrogen Bonding Inside Anionic Polymeric Brush Layer: Machine Learning-Driven Exploration of the Relative Roles of the Polymer Steric Effect, Charging, and Type of Screening Counterions

This paper employs a combination of all-atom molecular dynamics (MD) simulations and unsupervised machine learning (ML) for studying the water-water hydrogen bonds (HBs) inside the anionic poly-acrylic acid (PAA) brushes modeled using all-atom MD simulations. PAA brush layer with different charge fraction (f), namely f=0, f=0.25, and f=1, is considered. Water-water interactions, both inside and outside the brush layer, are represented through distinct clusters of tupules of variables representing distances associated with the interacting water molecules. While clusters representing the HBs are present for water inside and outside the brushes, several clusters representing the long-range water-water interactions are missing for the water molecules inside the highly charged (f=1) PAA brushes. More importantly, inside highly charged brushes, the edge of the clusters representing the water-water HBs is progressively shortened, as compared to that in the bulk. Both these results stem from the presence of the PAA brushes imparting the steric effect and the charge effect, or the effect associated with enhanced interactions of water molecules with PE charges and counterions, thereby disrupting the water connectivity. This water-charged-species interaction also increases the water-water HB angle, i.e., makes the water-water HBs less stable inside the highly charged PAA brush layer. The narrowing of the clusters representing the HBs and the alteration of the angle characterizing the HBs confirm that the conditions defining the water-water HBs change inside the PAA brush layer as a function of the charges on the PAA brush layer. Furthermore, we show that the use of the generic definition of HBs, as compared to using our simulation-motivated modified definition of water-water HBs, overpredict the number of water-water HBs inside the PAA brush layer. Finally, we employ this all-atom-MD-ML framework to quantify the effect of other types of screening counterions (Li + , Ca 2+ , and Y 3+ ions) in determining the water-water interactions and water-water HB properties inside the PAA brush layer. Furthermore, the findings of the present study, confirming the weakening of water-water HBs inside the PAA brush layer, points to the possibility that the water molecules will be more available for hydrating the brush layer and counterions, thereby leading to a more pronounced wetting of the PAA brush layer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Objective Phenotyping of Root System Architecture Using Image Augmentation and Machine Learning in Alfalfa (Medicago sativa L.)

Active breeding programs specifically for root system architecture (RSA) phenotypes remain rare; however, breeding for branch and taproot types in the perennial crop alfalfa is ongoing. Phenotyping in this and other crops for active RSA breeding has mostly used visual scoring of specific traits or subjective classification into different root types. While image-based methods have been developed, translation to applied breeding is limited. This research is aimed at developing and comparing image-based RSA phenotyping methods using machine and deep learning algorithms for objective classification of 617 root images from mature alfalfa plants collected from the field to support the ongoing breeding efforts. Our results show that unsupervised machine learning tends to incorrectly classify roots into a normal distribution with most lines predicted as the intermediate root type. Encouragingly, random forest and TensorFlow-based neural networks can classify the root types into branch-type, taproot-type, and an intermediate taproot-branch type with 86% accuracy. With image augmentation, the prediction accuracy was improved to 97%. Coupling the predicted root type with its prediction probability will give breeders a confidence level for better decisions to advance the best and exclude the worst lines from their breeding program. This machine and deep learning approach enables accurate classification of the RSA phenotypes for genomic breeding of climate-resilient alfalfa.

59 BASIC BIOLOGICAL SCIENCES↗

General-Purpose Unsupervised Cyber Anomaly Detection via Non-Negative Tensor Factorization

Distinguishing malicious anomalous activities from unusual but benign activities is a fundamental challenge for cyber defenders. Prior studies have shown that statistical user behavior analysis yields accurate detections by learning behavior profiles from observed user activity. These unsupervised models are able to generalize to unseen types of attacks by detecting deviations from normal behavior, without knowledge of specific attack signatures. However, approaches proposed to date based on probabilistic matrix factorization are limited by the information conveyed in a two-dimensional space. Non-negative tensor factorization, on the other hand, is a powerful unsupervised machine learning method that naturally models multi-dimensional data, capturing complex and multi-faceted details of behavior profiles. Herein, our new unsupervised statistical anomaly detection methodology matches or surpasses state-of-the-art supervised learning baselines across several challenging and diverse cyber application areas, including detection of compromised user credentials, botnets, spam e-mails, and fraudulent credit card transactions.

97 MATHEMATICS AND COMPUTING↗

Defect detection in atomic-resolution images via unsupervised learning with translational invariance

Abstract Crystallographic defects can now be routinely imaged at atomic resolution with aberration-corrected scanning transmission electron microscopy (STEM) at high speed, with the potential for vast volumes of data to be acquired in relatively short times or through autonomous experiments that can continue over very long periods. Automatic detection and classification of defects in the STEM images are needed in order to handle the data in an efficient way. However, like many other tasks related to object detection and identification in artificial intelligence, it is challenging to detect and identify defects from STEM images. Furthermore, it is difficult to deal with crystal structures that have many atoms and low symmetries. Previous methods used for defect detection and classification were based on supervised learning, which requires human-labeled data. In this work, we develop an approach for defect detection with unsupervised machine learning based on a one-class support vector machine (OCSVM). We introduce two schemes of image segmentation and data preprocessing, both of which involve taking the Patterson function of each segment as inputs. We demonstrate that this method can be applied to various defects, such as point and line defects in 2D materials and twin boundaries in 3D nanocrystals.

36 MATERIALS SCIENCE↗

Unraveling Hydrogen Induced Geochemical Reaction Mechanisms through Coupled Geochemical Modeling and Machine Learning

Underground hydrogen storage (UHS) provides a promising large-scale, long-term energy storage solution. A reasonable recovery of stored hydrogen is critical for a successful storage scheme. However, in subsurface reservoirs hydrogen is subject to active geochemical reactions that might result in hydrogen loss. In this study, we implemented a geochemical modeling approach coupled with an unsupervised machine learning technique called non-negative matrix factorization (NMF) to unravel the complex brine-rock-H 2 geochemical processes responsible for hydrogen losses, with particular focus on sulfate reduction reactions. NMF is applied to modeled mineral evolution and fluid component profiles to retrieve profiles that can be interpreted to more easily assess competing processes. NMF decouples simulated competing equilibrium reactions. This facilitates separation of overlapping reaction profiles from redox processes, dissolution fronts, and secondary precipitation while considering the effects of simulation parameters such as salinity, temperature, and total H 2 pressure. NMF successfully discriminates these competing effects in nonlinear ways, allowing robust interpretation. In addition, NMF reveals subtle coupled mineral associations and reaction fronts that are invisible to conventional model analysis. This integrated approach strengthens the conceptual understanding of complex nonlinear hydrogen-brine-rock interactions and advances geochemical research on UHS systems to resolve complexities in modeled geochemical systems without the need for direct experiments or prior knowledge. Furthermore, this study highlights the efficacy of combining geochemical modeling with machine learning techniques to enhance the interpretability of the intricate geochemical simulation output through deciphering the overlapping reaction path that cannot be achieved only using conventional analysis of geochemical models alone.

08 HYDROGEN↗

Spread spectrum time domain reflectometry (SSTDR) and frequency domain reflectometry (FDR) cable inspection using machine learning

Cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, justification for continued cable use must shift to a condition-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. The Pacific Northwest National Laboratory (PNNL) Accelerated and Real Time Experimental Nodal Analysis (ARENA) cable motor test bed was used to test the response of a commercial spread spectrum time domain reflectometry (SSTDR) system, a laboratory instrument software-controlled SSTDR, and a vector network analyzer-based frequency domain reflectometry (FDR) system to various cable anomalies. The three instrument systems were able to interrogate cables over a range of frequency bandwidths that can be helpful for human data analysis. Data were subjected to supervised and unsupervised machine learning (ML) analyses to distinguish normal undamaged cable responses from anomalous cable responses. Both supervised and unsupervised ML approaches produced encouraging results with an undamaged/anomalous prediction accuracy from 0.69% to 0.87%. Recommendations for further development and field implementation include increased and more balanced sample sets particularly including more training data.

SSTDR, FDR, Reflectometry, Machine Learning, ARENA↗

Comparison of Supervised and Un-Supervised Machine Learning Algorithms for Threat Detection and Scintillator Performance for Radiation Portal Monitoring

Following the events of September 11, 2001, international border crossing have been equipped with radiation portal monitors (RPMs) to identify illicit radioactive material. Polyvinyl toluene (PVT) scintillators are commonly used due to their low cost and reasonable maintainability, however they offer low spectral resolution. Despite the fact that over twenty years has transpired since this event, radioisotopes are still typically identified by hand-crafted classification algorithms, e.g., total counts or energy windowing, and exhibit relatively poor performance in detecting threats at the low false alarm rates required to support the stream of commerce. While some improvement to performance has been realized via the use of supervised machine learning, these classification algorithms typically utilize simulations in lieu of real data due to the sparsity of data for one or more classes. Accordingly, the performance of these algorithms is somewhat less than optimal when examining experiments or simulations with model mismatch. Consequently, in this work, we examine the application of a number of unsupervised machine learning, anomaly detection based algorithms, to circumvent the inverse crime when analyzing spectroscopy data for RPMs. We also compare anomaly detection results with those obtained via the use of supervised classification detection ML algorithms when model mismatch is introduced between the simulated threat items utilized for training/testing. Finally, we compared the performance of the PVT scintillators to those obtained with higher resolution detectors using both anomaly detection and supervised classification algorithms.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

Anomaly detection search for new resonances decaying into a Higgs boson and a generic new particle $X$ in hadronic final states using $\sqrt{s}$ = 13 TeV $pp$ collisions with the ATLAS detector

A search is presented for a heavy resonance $Y$ decaying into a Standard Model Higgs boson $H$ and a new particle $X$ in a fully hadronic final state. The full Large Hadron Collider run 2 dataset of proton-proton collisions at $\sqrt{s}$ = 13 TeV collected by the ATLAS detector from 2015 to 2018 is used and corresponds to an integrated luminosity of 139 fb –1 . The search targets the high $Y$-mass region, where the $H$ and $X$ have a significant Lorentz boost in the laboratory frame. A novel application of anomaly detection is used to define a general signal region, where events are selected solely because of their incompatibility with a learned background-only model. It is constructed using a jet-level tagger for signal-model-independent selection of the boosted $X$ particle, representing the first application of fully unsupervised machine learning to an ATLAS analysis. Two additional signal regions are implemented to target a benchmark $X$ decay into two quarks, covering topologies where the $X$ is reconstructed as either a single large-radius jet or two smallradius jets. The analysis selects Higgs boson decays into $b$$\overline{b}$, and a dedicated neural-network-based tagger provides sensitivity to the boosted heavy-flavor topology. No significant excess of data over the expected background is observed, and the results are presented as upper limits on the production cross section $σ$ ($pp$ → $Y$ → $XH$ → $q$$\overline{q}$$b$$\overline{b}$) for signals with $m_Y$ between 1.5 and 6 TeV and $m_X$ between 65 and 3000 GeV

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Nowcasting Earthquakes: Imaging the Earthquake Cycle in California With Machine Learning

We propose a new machine learning-based method for nowcasting earthquakes to image the time-dependent earthquake cycle. The result is a timeseries that may correspond to the process of stress accumulation and release. The timeseries are constructed by using principal component analysis of regional seismicity. The patterns are found as eigenvectors of the cross-correlation matrix of a collection of seismicity timeseries in a coarse grained regional spatial grid (pattern recognition via unsupervised machine learning). The eigenvalues of this matrix represent the relative importance of the various eigenpatterns. Using the eigenvectors and eigenvalues, we compute the weighted correlation timeseries of the regional seismicity. This timeseries has the property that the weighted correlation generally decreases prior to major earthquakes in the region, and increases suddenly just after a major earthquake occurs. As in a previous paper, we find that this method produces a nowcasting timeseries that resembles the hypothesized regional stress accumulation and release process characterizing the earthquake cycle. We then address the problem of whether the timeseries contain information regarding future large earthquakes. For this, we compute a receiver operating characteristic and determine the decision thresholds for several future time periods of interest (optimization via supervised machine learning). We find that signals can be detected that can be used to characterize the information content of the timeseries. These signals may be useful in assessing present and near-future seismic hazards.

58 GEOSCIENCES↗

Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) using Machine Learning

Modern retrofit construction practices use 3D point cloud data of the building envelope to obtain the as-built dimensions. However, manual segmentation by a trained professional is required to identify and measure window openings, door openings, and other architectural features, making the use of 3D point clouds labor-intensive. In this study, the Automatic point Cloud Building Envelope Segmentation (Auto-CuBES) algorithm is described, which can significantly reduce the time spent during point cloud segmentation. The Auto-CuBES algorithm inputs a 3D point cloud generated by commonly available surveying equipment and outputs a wire-frame model of the building envelope. Unsupervised machine learning methods were used to identify facades, windows, and doors while minimizing the number of calibration parameters. Additionally, Auto-CuBES generates a heat map of each facade indicating non-planar characteristics that are crucial for the optimization of connections used in overclad envelope retrofits. With a scan resolution of 3 mm, the resulting window dimensions showed a mean absolute error of 4.2 mm compared to manual laser measurements.

Maldonado Puente, Bryan↗

Confidentiality-preserving machine learning algorithms for soft-failure detection in optical communication networks

Automated fault management is at the forefront of next-generation optical communication networks. The increase in complexity of modern networks has triggered the need for programmable and software-driven architectures to support the operation of agile and self-managed systems. In these scenarios, the European Telecommunications Standards Institute zero-touch network and service management approach is imperative. The need for machine learning algorithms to process the large volume of telemetry data brings safety concerns as distributed cloud-computing solutions become the preferred approach for deploying reliable communication network automation. This paper’s contribution is twofold. First, we propose a simple yet effective method to guarantee the confidentiality of the telemetry data based on feature scrambling. The method allows the operation of third-party computational services without direct access to the full content of the collected data. Additionally, the effectiveness of four unsupervised machine learning algorithms for soft-failure detection is evaluated when applied to the scrambled telemetry data. The methods are based on factor analysis, principal component analysis, nonlinear principal component analysis, and singular value decomposition. Most dimensionality reduction algorithms have the common property that they can maintain similar levels of fault classification performance while hiding the data structure from unauthorized access. Evaluations of the proposed algorithms demonstrate this capability.

97 MATHEMATICS AND COMPUTING↗

Scaling Building Energy Audits through Machine Learning Methods on Novel Drone Image Data

Building energy audits are time-consuming and labor-intensive. This paper describes a new method using machine learning (ML) techniques on novel data sources (drone images) to improve the identification of building characteristics and retrofit opportunities, and thereby reduce the effort for audits. The new ML method includes: (1) Building footprint extraction using line extraction, polygonization, and polygon-merging, (2) Building envelope extraction using PIX4d modeling software to reconstruct a building 3D model, (3) Visualization tool for viewing images from the 3D model, (4) Window-to-wall ratio (WWR) using state-of-art deep neural network semantic segmentation, (5) Envelope thermal anomaly detection using an unsupervised machine learning clustering algorithm, and (6) Rooftop energy equipment detection based on an object detection algorithm. The testing of this method involved a comparison of additional ML-generated information overlaid on current ‘state-of-practice’ audit and remote assessment baselines using evaluation metrics: labor time and associated cost, marginal benefits of using ML-generated information in workflows for audits and remote assessments, integration potential with existing processes and tools, and replicability/scalability of the method. In two test buildings in California that had comprehensive drawings and meter data available, the ML method effectively generated a building footprint, envelope, rooftop equipment, WWR, and locations of envelope thermal anomalies. Projected target segments of the ML method are sites with minimal drawings and energy data, and underserved sectors such as multistoried housing, disadvantaged communities, and schools for which the ML method can enable identification of building asset characteristics and prioritization of envelope retrofits and decentralized energy equipment retrofits.

Singh, Reshma↗

Machine Learning of All Mycobacterium tuberculosis H37Rv RNA-seq Data Reveals a Structured Interplay between Metabolism, Stress Response, and Infection

Mycobacterium tuberculosis is one of the most consequential human bacterial pathogens, posing a serious challenge to 21st century medicine. A key feature of its pathogenicity is its ability to adapt its transcriptional response to environmental stresses through its transcriptional regulatory network (TRN). While many studies have sought to characterize specific portions of the M. tuberculosis TRN, and some studies have performed system-level analysis, few have been able to provide a network-based model of the TRN that also provides the relative shifts in transcriptional regulator activity triggered by changing environments. Here, we compiled a compendium of nearly 650 publicly available, high quality M. tuberculosis RNA-sequencing data sets and applied an unsupervised machine learning method to obtain a quantitative, top-down TRN. It consists of 80 independently modulated gene sets known as “iModulons,” 41 of which correspond to known regulons. These iModulons explain 61% of the variance in the organism’s transcriptional response. We show that iModulons (i) reveal the function of poorly characterized regulons, (ii) describe the transcriptional shifts that occur during environmental changes such as shifting carbon sources, oxidative stress, and infection events, and (iii) identify intrinsic clusters of regulons that link several important metabolic systems, including lipid, cholesterol, and sulfur metabolism. This transcriptome-wide analysis of the M. tuberculosis TRN informs future research on effective ways to study and manipulate its transcriptional regulation and presents a knowledge-enhanced database of all published high-quality RNA-seq data for this organism to date.

59 BASIC BIOLOGICAL SCIENCES↗

Distinguishing isotropic and anisotropic signals for X-ray total scattering using machine learning

Understanding structure–property relationships is essential for advancing technologies based on thin films. X-ray pair distribution function (PDF) analysis can access relevant atomic structure details spanning local-, mid- and long-range structure. While X-ray PDF has been adapted for thin films on amorphous substrates, measurements on single-crystal substrates are necessary to accurately determine structure origins for some thin film materials, especially those for which the substrate changes the accessible structure and properties. However, when measuring films on single-crystal substrates, high-intensity anisotropic Bragg spots saturate 2D detector images, overshadowing the thin films' isotropic scattering signal. This renders previous data processing methods for films on amorphous substrates unsuitable for films on single-crystal substrates. To address this measurement need, we developed IsoDAT2D, an innovative data processing approach using unsupervised machine learning algorithms. The program combines dimensionality reduction and clustering algorithms to separate thin film and single-crystal substrate X-ray scattering signals. We use SimDAT2D , a program we developed to generate simulated thin film data, to validate IsoDAT2D . Here we also use IsoDAT2D to isolate X-ray total scattering signal from a thin film on a single-crystal substrate. The resulting PDF data are compared with similar data processed using previous methods, especially substrate subtraction for single-crystal and amorphous substrates. PDF data from IsoDAT2D -identified X-ray total scattering data are significantly better than from single-crystal substrate subtraction, but not as reliable as PDF data from amorphous substrate subtraction. With IsoDAT2D , there are new opportunities to expand PDF to a wider variety of thin films, including those on single-crystal substrates, with which new structure–property relationships can be elucidated to enable fundamental understanding and technological advances.

36 MATERIALS SCIENCE↗

Investigation of acoustic waves under subsurface conditions to improve the predictions of rock mechanical properties and natural fracture characteristics

Mechanical properties and natural fracture characteristics are critical to investigate for subsurface engineering applications, including carbon storage, well drilling, and stimulation, as they govern rock stability, fluid flow, and mechanical behavior under stress. This dissertation integrates experimental and machine learning approaches to enhance the prediction and understanding of these properties by analyzing acoustic wave behavior under varied subsurface conditions. First, the influence of temperature, pore pressure, and supercritical CO2 (scCO2) saturation on poroelastic properties is examined using Gray Berea sandstone samples. The results show that temperature and pore pressure significantly affect the bulk modulus and Biot’s coefficient, while scCO2 saturation impacts rock compressibility, informing strategies for effective geological carbon storage. The study extends this understanding by experimentally evaluating the impact of reservoir depletion on the dynamic mechanical properties of the emerging Caney shale in South Oklahoma with the employment of unsupervised machine learning to predict static mechanical properties across the Caney shale. Integrating petrophysical data and chemostratigraphy, the workflow—featuring K-means clustering, principal component analysis (PCA), and inverse distance weighting (IDW)—improves stratigraphic characterization and the estimation of static-to-dynamic modulus ratios, which is vital for optimizing drilling and stimulation strategies. Finally, the work explores how natural fracture characteristics in shale influence acoustic waveforms and shear wave splitting (SWS) analysis. Experimental data on fractured samples under different stress and temperature conditions, combined with machine learning models such as K-nearest neighbors (KNN) and extreme gradient boosting (XGBoost), reveal key fracture properties impacting SWS and wave propagation. Together, these studies provide a comprehensive framework for linking acoustic wave behavior with rock properties, advancing the methods for monitoring and predicting geomechanical changes. The insights offered valuable implications for safer, more efficient CO2 injection, hydrocarbon extraction, and subsurface management.

Elkholy, Sherif↗

Machine learning enabled quantification of the hydrogen bonds inside the polyelectrolyte brush layer probed using all-atom molecular dynamics simulations

The configuration of densely grafted charged polyelectrolyte (PE) brushes is strongly dictated by the properties and behavior of the counterions that screen the PE brush charges and the solvent molecules (typically water) that solvate the brush molecules and these screening counterions. Only recently, efforts have been made to study the PE brushes atomistically, thereby shedding light on the properties of brush-supported ions and water molecules. However, even for such efforts, there are limitations associated with using a generic definition to estimate certain properties of water and ions inside the brush layer. For example, water–water hydrogen bonds (HBs) will behave differently for locations outside and inside the brush layer, given the fact that the densely closely grafted PE brush molecules create a soft nanoconfinement where the water connectivity becomes highly disrupted: therefore, using the same definition to quantify the HBs inside and outside the brush layer will be unwise. In this paper, we address this limitation by employing an unsupervised machine learning (ML) approach to predict the water–water hydrogen bonding inside a cationic PE brush layer modeled using all-atom molecular dynamics (MD) simulations. Here, the ML method, which relies on a clustering approach and uses the equilibrium coordinates of the water molecules (obtained from the all-atom MD simulations) as the input, is capable of identifying the structural modification of water–water HBs (revealed through appropriate clustering of the data) inside the PE brush layer induced soft nanoconfinement. Such capabilities would not have been possible by using a generic definition of the HBs. Our calculations lead to four key findings: (1) the clusters formed inside and outside the brush layer are structurally similar; (2) the margin of the cluster is shorter inside the PE brush layer confirming the possible disruption of the HBs inside the PE brush layer; (3) the average “hydrogen–acceptor-oxygen–donor-oxygen” angle that defines the HB is reduced for the HBs formed inside the brush layer; (4) the use of the generic definition (definition usable for characterizing the HBs in brush-free bulk) leads to an overprediction of the number of HBs formed inside the PE brush layer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗