A novel machine learning and deep learning semi-supervised approach for automatic detection of InSAR-based deformation hotspots
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
In this paper, we introduce an uplifted reduced order modeling (UROM) approach through the integration of standard projection based methods with long short-term memory (LSTM) embedding. Our approach has three modeling layers or components. In the first layer, we utilize an intrusive projection approach to model dynamics represented by the largest modes. The second layer consists of an LSTM model to account for residuals beyond this truncation. This closure layer refers to the process of including the residual effect of the discarded modes into the dynamics of the largest scales. However, the feasibility of generating a low rank approximation tails off for higher Kolmogorov n -width systems due to the underlying nonlinear processes. The third uplifting layer, called super-resolution, addresses this limited representation issue by expanding the span into a larger number of modes utilizing the versatility of LSTM. Therefore, our model integrates a physics-based projection model with a memory embedded LSTM closure and an LSTM based super-resolution model. In several applications, we exploit the use of Grassmann manifold to construct UROM for unseen conditions. We performed numerical experiments by using the Burgers and Navier-Stokes equations with quadratic nonlinearity. Finally, our results show robustness of the proposed approach in building reduced order models for parameterized systems and confirm the improved trade-off between accuracy and efficiency.
Understanding gene regulation is fundamental to plant improvement, but the lack of plant-specific single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) frameworks and cross-species databases has limited insights into cell-type-specific cellular regulation. Here we present ‘scPlantReg’, an integrated framework and database for plant scATAC-seq data. scPlantReg supports end-to-end analyses from raw data processing to biological interpretation and features ‘scATACtor’, a supervised machine-learning approach that outperforms existing tools for cell-type annotation. We applied scPlantReg to pearl millet to characterize cell-type-specific chromatin accessibility and identify validated activating and repressing accessible chromatin regions (ACRs), revealing WRKY transcription factors as potential regulators of xylem development. Furthermore, we reanalysed scATAC-seq datasets from 8 plant species, spanning 11 tissues and multiple developmental stages, enabling cross-species comparisons. Furthermore, these analyses uncovered conserved regulatory programmes, including AP2/EREBP-associated ACRs linked to cell wall development and cell-type-conserved TFs across grasses. Collectively, scPlantReg provides a general framework and resource for comparative regulatory analysis in plants.
Enhancers are important non-coding elements, but they have traditionally been hard to characterize experimentally. The development of massively parallel assays allows the characterization of large numbers of enhancers for the first time. Here, we developed a framework using Drosophila STARR-seq to create shape-matching filters based on meta-profiles of epigenetic features. We integrated these features with supervised machine-learning algorithms to predict enhancers. We further demonstrated that our model could be transferred to predict enhancers in mammals. We comprehensively validated the predictions using a combination of in vivo and in vitro approaches, involving transgenic assays in mice and transduction-based reporter assays in human cell lines (153 enhancers in total). The results confirmed that our model can accurately predict enhancers in different species without re-parameterization. Finally, we examined the transcription factor binding patterns at predicted enhancers versus promoters. Here, we demonstrated that these patterns enable the construction of a secondary model that effectively distinguishes enhancers and promoters.
Not Available
Precise neutrino energy reconstruction is essential for next-generation long-baseline oscillation experiments, yet current methods remain limited by large uncertainties in neutrino-nucleus interaction modeling. Even so, it is well established that different interaction channels produce systematically varying amounts of missing energy and therefore yield different reconstruction performance–information that standard calorimetric approaches do not exploit. We introduce a strategy that incorporates this structure by classifying events according to their underlying interaction type prior to energy reconstruction. Using supervised machine-learning techniques trained on labeled generator events, we leverage intrinsic kinematic differences among quasielastic scattering, meson-exchange current, resonance production, and deep-inelastic scattering processes. A cross-generator testing framework demonstrates that this classification approach is robust to microphysics mismodeling and, when applied to a simulated DUNE 𝜈 𝜇 disappearance analysis, yields improved accuracy and sensitivity at the 10%–20% level. These results highlight a practical path toward reducing reconstruction-driven systematics in future oscillation measurements.
M dwarfs are the most common type of star in the Galaxy, and because of their small size are favored targets for searches of Earth-sized transiting exoplanets. Current and upcoming all-sky spectroscopic surveys, such as the Large Sky Area Multi Fiber Spectroscopic Telescope (LAMOST), offer an opportunity to systematically determine physical properties of many more M dwarfs than has been previously possible. Here, we present new effective temperatures, radii, masses, and luminosities for 29,678 M dwarfs with spectral types M0–M6 in the first data release (DR1) of LAMOST. We derived these parameters from the supervised machine-learning code, The Cannon, trained with 1388 M dwarfs in the Transiting Exoplanet Survey Satellite Cool Dwarf Catalog that were also present in LAMOST with high signal-to-noise ratio (>250) spectra. Our validation tests show that the output parameter uncertainties are strongly correlated with the signal-to-noise of the LAMOST spectra, and we achieve typical uncertainties of 110 K in T{sub eff} (∼3%), 0.065 R{sub ⊙} (∼14%) in radius, 0.054 M{sub ⊙} (∼12%) in mass, and 0.012 L{sub ⊙} (∼20%) in luminosity. The model presented here can be rapidly applied to future LAMOST data releases, significantly extending the samples of well-characterized M dwarfs across the sky using new and exclusively data-based modeling methods.
Warm Neptunes offer a rich opportunity for understanding exo-atmospheric chemistry. With the upcoming James Webb Space Telescope (JWST), there is a need to elucidate the balance between investments in telescope time versus scientific yield. We use the supervised machine-learning method of the random forest to perform an information content (IC) analysis on a 11-parameter model of transmission spectra from the various NIRSpec modes. The three bluest medium-resolution NIRSpec modes (0.7–1.27 μm, 0.97–1.84 μm, 1.66–3.07 μm) are insensitive to the presence of CO. The reddest medium-resolution mode (2.87–5.10 μm) is sensitive to all of the molecules assumed in our model: CO, CO{sub 2}, CH{sub 4}, C{sub 2}H{sub 2}, H{sub 2}O, HCN, and NH{sub 3}. It competes effectively with the three bluest modes on the information encoded on cloud abundance and particle size. It is also competitive with the low-resolution prism mode (0.6–5.3 μm) on the inference of every parameter except for the temperature and ammonia abundance. We recommend astronomers to use the reddest medium-resolution NIRSpec mode for studying the atmospheric chemistry of 800–1200 K warm Neptunes; its corresponding high-resolution counterpart offers diminishing returns. We compare our findings to previous JWST IC analyses that favor the blue orders and suggest that the reliance on chemical equilibrium could lead to biased outcomes if this assumption does not apply. A simple, pressure-independent diagnostic for identifying chemical disequilibrium is proposed based on measuring the abundances of H{sub 2}O, CO, and CO{sub 2}.
We present a sample of 254,882 luminous red giant branch (LRGB) stars selected from the APOGEE and LAMOST surveys. By combining photometric and astrometric information from the Two Micron All Sky Survey and Gaia survey, the precise distances of the sample stars are determined by a supervised machine-learning algorithm: the gradient-boosted decision trees. To test the accuracy of the derived distances, member stars of globular clusters (GCs) and open clusters are used. The tests by cluster member stars show a precision of about 10% with negligible zero-point offsets, for the derived distances of our sample stars. The final sample covers a large volume of the Galactic disk(s) and halo of 0 < R < 30 kpc and |Z| ≤ 15 kpc. The rotation curve (RC) of the Milky Way across the radius of 5 ≲ R ≲ 25 kpc has been accurately measured with ~54,000 stars of the thin disk population selected from the LRGB sample. The derived RC shows a weak decline along R with a gradient of -1.83 ± 0.02 (stat.) ± 0.07 (sys.) km s -1 kpc -1 , in excellent agreement with the results measured by previous studies. The circular velocity at the solar position, yielded by our RC is 234.04 ± 0.08 (stat.) ± 1.36 (sys.) km s -1 , again in great consistency with other independent determinations. From the newly constructed RC, as well as constraints from other data, we have constructed a mass model for our Galaxy, yielding a mass of the dark matter halo of M 200 = (8.05 ± 1.15) × 10 11 M ⊙ with a corresponding radius of R 200 = 192.37 ± 9.24 kpc and a local dark matter density of 0.39 ± 0.03 GeV cm -3 .
Machine-learning methods and apparatus are provided to solve blind source separation problems with an unknown number of sources and having a signal propagation model with features such as wave-like propagation, medium-dependent velocity, attenuation, diffusion, and/or advection, between sources and sensors. In exemplary embodiments, multiple trials of non-negative matrix factorization are performed for a fixed number of sources, with selection criteria applied to determine successful trials. A semi-supervised clustering procedure is applied to trial results, and the clustering results are evaluated for robustness using measures for reconstruction quality and cluster separation. The number of sources is determined by comparing these measures for different trial numbers of sources. Source locations and parameters of the signal propagation model can also be determined. Disclosed methods are applicable to a wide range of spatial problems including chemical dispersal, pressure transients, and electromagnetic signals, and also to non-spatial problems such as cancer mutation.
Machine-learning methods and apparatus are provided to solve blind source separation problems with an unknown number of sources and having a signal propagation model with features such as wave-like propagation, medium-dependent velocity, attenuation, diffusion, and/or advection, between sources and sensors. In exemplary embodiments, multiple trials of non-negative matrix factorization are performed for a fixed number of sources, with selection criteria applied to determine successful trials. A semi-supervised clustering procedure is applied to trial results, and the clustering results are evaluated for robustness using measures for reconstruction quality and cluster separation. The number of sources is determined by comparing these measures for different trial numbers of sources. Source locations and parameters of the signal propagation model can also be determined. Disclosed methods are applicable to a wide range of spatial problems including chemical dispersal, pressure transients, and electromagnetic signals, and also to non-spatial problems such as cancer mutation.
Neutrinoless double-beta decay ($0\nu\beta\beta$) is a rare nuclear process that, if observed, will provide insight into the nature of neutrinos and help explain the matter-antimatter asymmetry in the Universe. The large enriched germanium experiment for neutrinoless double-beta decay (LEGEND) will operate in two phases to search for $0\nu\beta\beta$. The first (second) stage will employ 200 (1000) kg of High-Purity Germanium (HPGe) enriched in 76 Ge to achieve a half-life sensitivity of 10 27 (10 28 ) years. In this study, we present a semi-supervised data-driven approach to remove non-physical events captured by HPGe detectors powered by a novel artificial intelligence model. We utilize affinity propagation to cluster waveform signals based on their shape and a support vector machine to classify them into different categories. We train, optimize, and test our model on data taken from a natural abundance HPGe detector installed in the Full Chain Test experimental stand at the University of North Carolina at Chapel Hill. We demonstrate that our model yields a maximum sacrifice of physics events of $0.024 ^{+0.004}_{-0.003} \%$ after data cleaning. Our model is being used to accelerate data cleaning development for LEGEND-200 and will serve to improve data cleaning procedures for LEGEND-1000.
Fusion power production in tokamaks uses discharge configurations that risk producing strong type I edge localized modes. The largest of these modes will likely increase impurities in the plasma and potentially damage plasma facing components, such as the protective heat and particle divertor. Machine learning-based prediction and control may provide for the automatic detection and mitigation of these damaging modes before they grow too large to suppress. To that end, large labeled datasets are required for the supervised training of machine learning models. We present an algorithm that achieves 97.7% precision when automatically labeling edge localized modes in the large DIII-D tokamak discharge database. The algorithm has no user controlled parameters and is largely robust to tokamak and plasma configuration changes. This automatically labeled database of events can subsequently feed future training of machine learning models aimed at autonomous edge localized mode control and suppression.
Cables are initially qualified for nuclear power plant use for 40 years. As plants extend their operating license to 60 and 80 years, justification for continued cable use must shift to a condition-based approach since it is cost prohibitive to completely replace cables that are likely still capable of performing their design function. The Pacific Northwest National Laboratory (PNNL) Accelerated and Real Time Experimental Nodal Analysis (ARENA) cable motor test bed was used to test the response of a commercial spread spectrum time domain reflectometry (SSTDR) system, a laboratory instrument software-controlled SSTDR, and a vector network analyzer-based frequency domain reflectometry (FDR) system to various cable anomalies. The three instrument systems were able to interrogate cables over a range of frequency bandwidths that can be helpful for human data analysis. Data were subjected to supervised and unsupervised machine learning (ML) analyses to distinguish normal undamaged cable responses from anomalous cable responses. Both supervised and unsupervised ML approaches produced encouraging results with an undamaged/anomalous prediction accuracy from 0.69% to 0.87%. Recommendations for further development and field implementation include increased and more balanced sample sets particularly including more training data.
A Physics-Informed Machine Learning (PIML) framework for better system vulnerability assessment and faster corrective action recommendation. This framework consists of deriving physics-informed priors and smart sampling algorithms to help reduce the data samples of grid models and the dimension of simulation outputs, yielding small yet representative subset of the complex system, and both supervised and unsupervised machine learning (ML) algorithms for designing corrective actions
Abandoned coal mine lands (AMLs) represent one of the most persistent environmental challenges in the United States. Prior to the enactment of the Surface Mining Control and Reclamation Act (SMCRA) in 1977, coal mining operations were not legally required to reclaim disturbed lands, leaving behind approximately 500,000 AML sites nationwide. These sites pose severe environmental and health risks, including acid mine drainage, soil and water contamination, and spontaneous combustion of waste coal piles. Millions of Americans live within one mile of these AMLs, underscoring the urgency of remediation. Traditional reclamation practices, such as planting cool-season grasses, often fail to fully restore ecological function or leverage the economic potential of these lands. This project addressed these challenges by developing integrated strategies for resource recovery, land reclamation, and sustainable energy production. This project evaluated an integrated strategy to convert this liability into an opportunity by recovering waste coal and co-firing it with switchgrass (Panicum virgatum L.) cultivated on reclaimed or marginal AML areas in existing coal-fired power plants. Switchgrass not only provides a renewable feedstock but also aids in land reclamation and carbon sequestration. 1) Remote Sensing and Machine Learning for Waste Coal Identification Using Sentinel-2 satellite imagery and supervised classification, we applied four machine learning models to detect historical waste coal piles. Random Forest achieved the highest accuracy (precision: 86%, recall: 77%). Time-series analysis revealed gradual vegetation recovery since 1986, indicating natural reclamation processes in historical sites, while active mining areas showed ongoing disturbance. This workflow enables scalable monitoring and prioritization of reclamation efforts. 2) UAS-Based Stockpile Volume Estimation To quantify recoverable waste coal, we evaluated Unmanned Aerial Systems (UAS) equipped with Light Detection and Ranging (LiDAR) and multispectral sensors. Structure-from-Motion (SfM) photogrammetry combined with interpolated Digital Terrain Models (DTMs) achieved strong agreement with LiDAR reference volumes (Root Mean Square Error (RMSE) ≈147 m 3 , Mean Absolute Percentage Error (MAPE) ≈2%). Sensitivity analysis confirmed that spatial resolution significantly influences accuracy, emphasizing the need for high-resolution data for precise volume estimation. This approach offers a scalable, cost-effective, and accurate alternative to conventional ground-based surveys. 3) Switchgrass Cultivation for Bioenergy and Water Quality Improvement We assessed the hydrological and water quality impacts of converting AMLs to switchgrass production areas using the Soil and Water Assessment Tool (SWAT). Results showed that converting 10% of the watershed area into the switchgrass production zone reduced streamflow by 3.1%, total suspended solids by 18.1%, total nitrogen by 7.6%, and total phosphorus by 6.2%, while achieving biomass yields of 8.6–9.2 metric tons per hectare. These findings highlight switchgrass as a dual-benefit strategy for land reclamation and bioenergy feedstock production. 4) Integrated Co-Firing and CCS for Carbon-Negative Power Generation We modeled co-firing scenarios using the Power Plant Flexible Model (PPFM) to evaluate plant efficiency, greenhouse gas (GHG) emissions, and levelized cost of electricity (LCOE). Without carbon capture and storage (CCS), increasing switchgrass co-firing ratios reduced LCOE from $\$$150/MWh at 0% biomass to $\$$110/MWh at full substitution. Under CCS, costs remained higher (~$\$$250/MWh at 0% biomass) but decreased to $\$$200/MWh at 100% biomass, while enabling net-zero or carbon-negative electricity due to switchgrass sequestration benefits. Although CCS introduced efficiency penalties, pairing it with biomass co-firing offset these impacts and maximized climate benefits. Overall, optimizing co-firing ratios between 60-100%, supported by reliable logistics and storage strategies, emerged as a practical pathway to balance affordability, sustainability, and net-zero or negative GHG emissions while promoting productive reuse of AMLs.
Machine learning (ML), including deep learning (DL), has become increasingly popular in the last few years due to its continually outstanding performance. In this context, we apply machine learning techniques to "learn" the microstructure using both supervised and unsupervised DL techniques. In particular, we focus (1) on the localization problem bridging (micro)structure (localized) property using supervised DL and (2) on the microstructure reconstruction problem in latent space using unsupervised DL. The goal of supervised and semi-supervised DL is to replace crystal plasticity finite element model (CPFEM) that maps from (micro)structure (localized) property, and implicitly the (micro)structure (homogenized) property relationships, while the goal of unsupervised DL is (1) to represent high-dimensional microstructure images in a non-linear low-dimensional manifold, and (2) to discover a way to interpolate microstructures via latent space associating with latent microstructure variables. At the heart of this report is the applications of several common DL architectures, including convolutional neural networks (CNN), autoencoder (AE), and generative adversarial network (GAN), to multiple microstructure datasets, and the quest of neural architecture search for optimal DL architectures.
The following topics are addressed in this seminar presentation: The author's background; What is leading-order analysis?; What are supervised and unsupervised machine learning and what is artificial intelligence?; The definition of AI; and, Conclusions and outlook.