Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Numerical characterization of support recovery in sparse regression with correlated design

Sparse regression is employed in diverse scientific settings as a feature selection method. A pervasive aspect of scientific data is the presence of correlations between predictive features. These correlations hamper both feature selection and estimation and jeopardize conclusions drawn from estimated models. On the other hand, theoretical results on sparsity-inducing regularized regression have largely addressed conditions for selection consistency via asymptotics, and disregard the problem of model selection, whereby regularization parameters are chosen. In this numerical study, we address these issues through exhaustive characterization of the performance of several regression estimators, coupled with a range of model selection strategies. These estimators and selection criteria were examined across correlated regression problems with varying degrees of signal to noise, distributions of non-zero model coefficients, and model sparsity. Our results reveal a fundamental tradeoff between false positive and false negative control in all regression estimators and model selection criteria examined. Additionally, we numerically explore a transition point modulated by the signal-to-noise ratio and spectral properties of the design covariance matrix at which the selection accuracy of all considered algorithms degrades. Overall, we find that SCAD coupled with BIC or empirical Bayes model selection performs the best feature selection across the regression problems considered.

97 MATHEMATICS AND COMPUTING↗

Sparse regression for plasma physics

Many scientific problems can be formulated as sparse regression, i.e., regression onto a set of parameters when there is a desire or expectation that some of the parameters are exactly zero or do not substantially contribute. This includes many problems in signal and image processing, system identification, optimization, and parameter estimation methods such as Gaussian process regression. Sparsity facilitates exploring high-dimensional spaces while finding parsimonious and interpretable solutions. In the present work, we illustrate some of the important ways in which sparse regression appears in plasma physics and point out recent contributions and remaining challenges to solving these problems in this field. Further, a brief review is provided for the optimization problem and the state-of-the-art solvers, especially for constrained and high-dimensional sparse regression.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Characterization of Acoustic Emissions From Analogue Rocks Using Sparse Regression‐DMDc

Abstract Moisture loss in rock is known to generate acoustic emissions (AE). Phenomena that result in AE during drying are related to the movement of fluids through the pores and induced‐cracks that arise from differential mineral shrinkage, especially in clay‐bearing rock. AE from the movement of fluids occurs from the reconfiguration of fluid interfaces during drying, while AE from mineral shrinkage involves the debonding within or between minerals. Here, analogue rock samples were used to examine the differences in the AE signatures when one or both AE source‐types are present. An unsupervised sparse regression model, Dynamic Mode Decomposition with control, that extends Dynamic Mode Decomposition is used to characterize the AE signals recorded during the drying of porous analogue rock samples fabricated with ordinary Portland cement, with and without clay. This method can effectively and accurately reconstruct acoustic signals emitted from samples that only experience moisture loss without cracking. However, the method struggles to reconstruct signals from samples with intricate crack networks that formed during drying because AE generating mechanisms can emit contemporaneously, and the resulting waves propagate through drying‐induced cracks that can lead to multiple internal reflections. Thus, the differential reconstruction accuracy of time series generated by different underlying physical processes provides a robust filter for reducing large data catalogs. In general, both dynamics and sparse initiating events are learned directly from data and this method exposes a data hierarchy based on the complexity of the intrinsic dynamics.

58 GEOSCIENCES↗

Degeneracy engineering for classical and quantum annealing: A case study of sparse linear regression in collider physics

Classical and quantum annealing are computing paradigms that have been proposed to solve a wide range of optimization problems. In this paper, we aim to enhance the performance of annealing algorithms by introducing the technique of degeneracy engineering, through which the relative degeneracy of the ground state is increased by modifying a subset of terms in the objective Hamiltonian. We illustrate this novel approach by applying it to the example of ℓ 0 -norm regularization for sparse linear regression, which is, in general, an NP-hard optimization problem. Specifically, we show how to cast ℓ 0 -norm regularization as a quadratic unconstrained binary optimization (QUBO) problem, suitable for implementation on annealing platforms. As a case study, we apply this QUBO formulation to energy flow polynomials in high-energy collider physics, finding that degeneracy engineering substantially improves the annealing performance. Furthermore, our results motivate the application of degeneracy engineering to a variety of regularized optimization problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Clinical knowledge extraction via sparse embedding regression (KESER) with multi-center large scale electronic health record data

The increasing availability of electronic health record (EHR) systems has created enormous potential for translational research. However, it is difficult to know all the relevant codes related to a phenotype due to the large number of codes available. Traditional data mining approaches often require the use of patient-level data, which hinders the ability to share data across institutions. In this project, we demonstrate that multi-center large-scale code embeddings can be used to efficiently identify relevant features related to a disease of interest. We constructed large-scale code embeddings for a wide range of codified concepts from EHRs from two large medical centers. We developed knowledge extraction via sparse embedding regression (KESER) for feature selection and integrative network analysis. We evaluated the quality of the code embeddings and assessed the performance of KESER in feature selection for eight diseases. Besides, we developed an integrated clinical knowledge map combining embedding data from both institutions. The features selected by KESER were comprehensive compared to lists of codified data generated by domain experts. Features identified via KESER resulted in comparable performance to those built upon features selected manually or with patient-level data. The knowledge map created using an integrative analysis identified disease-disease and disease-drug pairs more accurately compared to those identified using single institution data. Analysis of code embeddings via KESER can effectively reveal clinical knowledge and infer relatedness among codified concepts. KESER bypasses the need for patient-level data in individual analyses providing a significant advance in enabling multi-center studies using EHR data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Zero-truncated Poisson regression for sparse multiway count data corrupted by false zeros

Abstract We propose a novel statistical inference methodology for multiway count data that is corrupted by false zeros that are indistinguishable from true zero counts. Our approach consists of zero-truncating the Poisson distribution to neglect all zero values. This simple truncated approach dispenses with the need to distinguish between true and false zero counts and reduces the amount of data to be processed. Inference is accomplished via tensor completion that imposes low-rank tensor structure on the Poisson parameter space. Our main result shows that an $N$-way rank-$R$ parametric tensor $\boldsymbol{\mathscr{M}}\in (0,\infty )^{I\times \cdots \times I}$ generating Poisson observations can be accurately estimated by zero-truncated Poisson regression from approximately $IR^2\log _2^2(I)$ non-zero counts under the nonnegative canonical polyadic decomposition. Our result also quantifies the error made by zero-truncating the Poisson distribution when the parameter is uniformly bounded from below. Therefore, under a low-rank multiparameter model, we propose an implementable approach guaranteed to achieve accurate regression in under-determined scenarios with substantial corruption by false zeros. Several numerical experiments are presented to explore the theoretical results.

97 MATHEMATICS AND COMPUTING↗

Stochastic AC optimal power flow: A data-driven approach

There is an emerging need for efficient solutions to stochastic AC Optimal Power Flow (AC-OPF) to ensure optimal and reliable grid operations in the presence of increasing demand and generation uncertainty. Herein this paper presents a highly scalable data-driven algorithm for stochastic AC-OPF that has extremely low sample requirement. The novelty behind the algorithm’s performance involves an iterative scenario design approach that merges information regarding constraint violations in the system with data-driven sparse regression. Compared to conventional methods with random scenario sampling, our approach is able to provide feasible operating points for realistic systems with much lower sample requirements. Furthermore, multiple sub-tasks in our approach can be easily paralleled and based on historical data to enhance its performance and application. We demonstrate the computational improvements of our approach through simulations on different test cases in the IEEE PES PGLib-OPF benchmark library.

42 ENGINEERING↗

Greedy permanent magnet optimization

Abstract A number of scientific fields rely on placing permanent magnets in order to produce a desired magnetic field. We have shown in recent work that the placement process can be formulated as sparse regression. However, binary, grid-aligned solutions are desired for realistic engineering designs. We now show that the binary permanent magnet problem can be formulated as a quadratic program with quadratic equality constraints, the binary, grid-aligned problem is equivalent to the quadratic knapsack problem with multiple knapsack constraints, and the single-orientation-only problem is equivalent to the unconstrained quadratic binary problem. We then provide a set of simple greedy algorithms for solving variants of permanent magnet optimization, and demonstrate their capabilities by designing magnets for stellarator plasmas. The algorithms can a-priori produce sparse, grid-aligned, binary solutions. Despite its simple design and greedy nature, we provide an algorithm that compares with or even outperforms the state-of-the-art algorithms while being substantially faster, more flexible, and easier to use.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Extreme sparsification of physics-augmented neural networks for interpretable model discovery in mechanics

Data-driven constitutive modeling with neural networks has received increased interest in recent years due to its ability to easily incorporate physical and mechanistic constraints and to overcome the challenging and time-consuming task of formulating phenomenological constitutive laws that can accurately capture the observed material response. However, even though neural network-based constitutive laws have been shown to generalize proficiently, the generated representations are not easily interpretable due to their high number of trainable parameters. Sparse regression approaches exist that allow for obtaining interpretable expressions, but the user is tasked with creating a library of model forms which by construction limits their expressiveness to the functional forms provided in the libraries. Here, in this work, we propose to train regularized physics-augmented neural network-based constitutive models utilizing a smoothed version of $L^0$-regularization. This aims to maintain the trustworthiness inherited by the physical constraints, but also enables interpretability which has not been possible thus far on any type of machine learning-based constitutive model where model forms were not assumed a priori but were actually discovered. During the training process, the network simultaneously fits the training data and penalizes the number of active parameters, while also ensuring constitutive constraints such as thermodynamic consistency. We show that the method can reliably obtain interpretable and trustworthy constitutive models for compressible and incompressible hyperelasticity, yield functions, and hardening models for elastoplasticity, using synthetic and experimental data. This work aims to set a new paradigm for interpretable machine learning models in the broad area of solid mechanics where low and limited data is available along with prior knowledge of physical constraints that the learned maps need to obey. This paradigm can potentially be extended to a broader spectrum of scientific exploration.

Data-driven constitutive models↗

Permutation-adapted complete and independent basis for atomic cluster expansion descriptors

Atomic cluster expansion (ACE) methods provide a systematic way to describe particle local environments of arbitrary body order. For practical applications it is often required that the basis of cluster functions be symmetrized with respect to rotations and permutations. Existing methodologies yield sets of symmetrized functions that are over-complete. These methodologies thus require an additional numerical procedure, such as singular value decomposition (SVD), to eliminate redundant functions. In this work, it is shown that analytical linear relationships for subsets of cluster functions may be derived using recursion and permutation properties of generalized Wigner symbols. From these relationships, subsets (blocks) of cluster functions can be selected such that, within each block, functions are guaranteed to be linearly independent. It is conjectured that this block-wise independent set of permutation-adapted rotation and permutation invariant (PA-RPI) functions forms a complete, independent basis for ACE. Along with the first analytical proofs of block-wise linear dependence of ACE cluster functions and other theoretical arguments, numerical results are offered to demonstrate this. The utility of the method is demonstrated in the development of an ACE interatomic potential for tantalum. Using the new basis functions in combination with Bayesian compressive sensing sparse regression, some high degree descriptors are observed to persist and help achieve high-accuracy models.

Angular momentum↗

Discovering equations that govern experimental materials stability under environmental stress using scientific machine learning

Abstract While machine learning (ML) in experimental research has demonstrated impressive predictive capabilities, extracting fungible knowledge representations from experimental data remains an elusive task. In this manuscript, we use ML to infer the underlying differential equation (DE) from experimental data of degrading organic-inorganic methylammonium lead iodide (MAPI) perovskite thin films under environmental stressors (elevated temperature, humidity, and light). Using a sparse regression algorithm, we find that the underlying DE governing MAPI degradation across a broad temperature range of 35 to 85 °C is described minimally by a second-order polynomial. This DE corresponds to the Verhulst logistic function, which describes reaction kinetics analogous to self-propagating reactions. We examine the robustness of our conclusions to experimental variance and Gaussian noise and describe the experimental limits within which this methodology can be applied. Our study highlights the promise and challenges associated with ML-aided scientific discovery by demonstrating its application in experimental chemical and materials systems.

36 MATERIALS SCIENCE↗

Data-Driven Discovery of Active Nematic Hydrodynamics

Active nematics can be modeled using phenomenological continuum theories that account for the dynamics of the nematic director and fluid velocity through partial differential equations (PDEs). While these models provide a statistical description of the experiments, the relevant terms in the PDEs and their parameters are usually identified indirectly. We adapt a recently developed method to automatically identify optimal continuum models for active nematics directly from spatio-temporal data, via sparse regression of the coarse-grained fields onto generic low order PDEs. After extensive benchmarking, we apply the method to experiments with microtubule-based active nematics, finding a surprisingly minimal description of the system. Furthermore, our approach can be generalized to gain insights into active gels, microswimmers, and diverse other experimental active matter systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Modeling Intercalation Chemistry with Multiredox Reactions by Sparse Lattice Models in Disordered Rocksalt Cathodes

Modern battery materials can contain many elements with substantial site disorder, and their configurational state has been shown to be critical for their performance. The intercalation voltage profile is a critical parameter to evaluate the performance of energy storage. The application of commonly used cluster expansion techniques to model the intercalation thermodynamics of such systems ab initio is challenged by the combinatorial increase in configurational degrees of freedom as the number of species grows. Such challenges necessitate the efficient generation of lattice models without overfitting and proper sampling of the configurational space under the requirement of charge balance in ionic systems. In this work, we introduce a combined approach that addresses these challenges by (1) constructing a robust cluster expansion Hamiltonian using the sparse regression technique, including -norm regularization and structural hierarchy; and (2) implementing semigrand-canonical Monte Carlo to sample charge-balanced ionic configurations using the table-exchange method and an ensemble average approach. These techniques are applied to a disordered rocksalt oxyfluoride (LMNOF) that is part of a family of promising earth-abundant cathode materials. The simulated voltage profile is found to be in good agreement with experimental data and particularly provides a clear demonstration of the and oxygen contributions to the redox potential as a function of content.

25 ENERGY STORAGE↗

Role of physics in physics-informed machine learning

Physical systems are characterized by inherent symmetries, one of which is encapsulated in the units of their parameters and system states. These symmetries enable a lossless order-reduction, e.g., via dimensional analysis based on the Buckingham theorem. Despite the latter's benefits, machine learning (ML) strategies for the discovery of constitutive laws seldom subject experimental and/or numerical data to dimensional analysis. We demonstrate the potential of dimensional analysis to significantly enhance the interpretability and generalizability of ML-discovered secondary laws. Our numerical experiments with creeping fluid flow past solid ellipsoids show how dimensional analysis enable both deep neural networks and sparse regression reproduce old results, e.g., Stokes law for a sphere, and generate new ones, e.g., an expression for an ellipsoid misaligned with the flow direction. Furthermore, our results suggest the need to incorporate other physics-based symmetries and invariances into ML-based techniques for equation discovery.

97 MATHEMATICS AND COMPUTING↗

Chemical mixture exposure patterns and obesity among U.S. adults in NHANES 2005–2012

The effect of chemical exposure on obesity has raised great concerns. Real-world chemical exposure always imposes mixture impacts, however their exposure patterns and the corresponding associations with obesity have not been fully evaluated. To discover obesity-related mixed chemical exposure patterns in the general U.S. population. Sparse Decompositional Regression (SDR), a model adapted from sparse representation learning technique, was developed to identify exposure patterns of chemical mixtures with exclusion (non-targeted model) and inclusion (targeted model) of health outcomes. We assessed the relationships between the identified chemical mixture patterns and obesity-related indexes. We also conducted a comprehensive evaluation of this SDR model by comparing to the existing models, including generalized linear regression model (GLM), principal component analysis (PCA), and Bayesian kernel machine regression (BKMR). Eight core exposure patterns were identified using the non-targeted SDR model. Patterns of high levels of MEP, high levels of naphthalene metabolites (ΣOH-Nap), and a pattern of high exposure levels of MCOP, MCNP, and MCPP were positively associated with obesity. Patterns of high levels of BP3, and a pattern of higher mixed levels of MPB, PPB, and MEP were found to have negative associations. Associations were strengthened using the targeted SDR model. In the single chemical analysis by GLM, BP3, MBP, PPB, MCOP, and MCNP showed significant associations with obesity or body indexes. The SDR model exceeded the performance of PCA in pattern identification. Both SDR and BKMR identified a positive contribution of ΣOH-Nap and MCOP, as well as a negative contribution of BP3 and PPB to obesity. Our study identified five core exposure patterns of chemical mixtures significantly associated with obesity using the newly developed SDR model. The SDR model could open a new avenue for assessing health effects of environmental mixture contaminants.

54 ENVIRONMENTAL SCIENCES↗