Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Chemical identification of new particle formation and growth precursors through positive matrix factorization of ambient ion measurements

Abstract. In the lower troposphere, rapid collisions between ions and trace gases result in the transfer of positive charge to the highest proton affinity species and negative charge to the lowest proton affinity species. Measurements of the chemical composition of ambient ions thus provide direct insight into the most acidic and basic trace gases and their ion–molecule clusters – compounds thought to be important for new particle formation and growth. We deployed an atmospheric pressure interface time-of-flight mass spectrometer (APi-ToF) to measure ambient ion chemical composition during the 2016 Holistic Interactions of Shallow Clouds, Aerosols, and Land Ecosystems (HI-SCALE) campaign at the United States Department of Energy Atmospheric Radiation Measurement facility in the Southern Great Plains (SGP), an agricultural region. Cations and anions were measured for alternating periods of ∼ 24 h over 1 month. We use binned positive matrix factorization (binPMF) and generalized Kendrick analysis (GKA) to obtain information about the chemical formulas and temporal variation in ionic composition without the need for averaging over a long timescale or a priori high-resolution peak fitting. Negative ions consist of strong acids including sulfuric and nitric acid, organosulfates, and clusters of NO3- with highly oxygenated organic molecules (HOMs) derived from monoterpene (MT) and sesquiterpene (SQT) oxidation. Organonitrates derived from SQTs account for most of the HOM signal. Combined with the diel profiles and back trajectory analysis, these results suggest that NO3 radical chemistry is active at this site. SQT oxidation products likely contribute to particle growth at the SGP site. The positive ions consist of bases including alkyl pyridines and amines and a series of high-mass species. Nearly all the positive ions contained only one nitrogen atom and in general support ammonia and amines as being the dominant bases that could participate in new particle formation. Overall, this work demonstrates how APi-ToF measurements combined with binPMF analysis can provide insight into the temporal evolution of compounds important for new particle formation and growth.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluation of data driven low-rank matrix factorization for accelerated solutions of the Vlasov equation

Low-rank methods have shown success in accelerating simulations of a collisionless plasma described by the Vlasov equation, but still rely on computationally costly linear algebra every time step. We propose a data-driven factorization method using artificial neural networks, specifically with convolutional layer architecture, that trains on existing simulation data. At inference time, the model outputs a low-rank decomposition of the distribution field of the charged particles, and we demonstrate that this step is faster than the standard linear algebra technique. Numerical experiments show that the method achieves comparable reconstruction accuracy for interpolation tasks, generalizing to unseen test data in a manner beyond just memorizing training data; patterns in factorization also inherently followed the same numerical trend as those within algebraic methods (e.g., truncated singular-value decomposition). However, when training on the first 70% of a time-series data and testing on the remaining 30%, the method fails to meaningfully extrapolate. Despite this limiting result, the technique may have benefits for simulations in a statistical steady-state or otherwise showing temporal stability. These results suggest that while the model offers a computationally efficient alternative for datasets with temporal stability, its current formulation is best suited for interpolation rather than for predicting future states in time-evolving systems. This study thus lays the groundwork for further refinement of neural network-based approaches to low-rank matrix factorization in high-dimensional plasma simulations.

97 MATHEMATICS AND COMPUTING↗

Seasonal Disorder in Urban Traffic Patterns: A Low Rank Analysis

This article proposes several advances to sparse nonnegative matrix factorization (SNMF) as a way to identify large-scale patterns in urban traffic data. The input to our model is traffic counts organized by time and location. Nonnegative matrix factorization additively decomposes this information, organized as a matrix, into a linear sum of temporal signatures. Penalty terms encourage this factorization to concentrate on only a few temporal signatures, with weights which are not too large. Our interest here is to quantify and compare the regularity of traffic behavior, particularly across different broad temporal windows. In addition to the rank and error, we adapt a measure introduced by Hoyer to quantify sparsity in the representation. Combining these, we construct several curves which quantify error as a function of rank (the number of possible signatures) and sparsity; as rank goes up and sparsity goes down, the approximation can be better and the error should decreases. Plots of several such curves corresponding to different time windows leads to a way to compare disorder/order at different time scalewindows. In this paper, we apply our algorithms and procedures to study a taxi traffic dataset from New York City. In this dataset, we find weekly periodicity in the signatures, which allows us an extra framework for identifying outliers as significant deviations from weekly medians. We then apply our seasonal disorder analysis to the New York City traffic data and seasonal (spring, summer, winter, fall) time windows. We do find seasonal differences in traffic order.

97 MATHEMATICS AND COMPUTING↗

Clustering High-dimensional Toxicogenomics Data with Rare Signals

Toxicogenomics studies the gene and protein activities to drug treatments or toxic exposures. As the drugs and genes are numerous, toxicogenomics data are naturally high dimensional, with dimension sizes up to millions. In addition, the distribution of toxicogenomics data is oftentimes skewed, and they contain rare but important signals representing a cell or organism’s response to toxicity. The combination of high dimension and extremely skewed distribution of toxicogenomics data makes clustering analysis extremely challenging.We present our study of clustering toxicogenomics data using classical approaches such as principal component analysis as well as deep learning approaches such as auto-encoders. Our experiments show that these approaches fail to preserve rare signals and produce high-quality clusters. We then explore augmenting matrix factorization with deep learning techniques such as attention mechanism to produce latent representations for clustering. Our technique is able to better preserve rare signals after dimensionality reduction than prior approaches. Furthermore, we combine our augmented matrix factorization with a mechanism similar to autoencoder to balance separable clusters and low regeneration errors. Our experiments demonstrate better clustering with our proposed approach.

Cong, Guojing↗

Multiscale Reactive Model for 1,3,5-Triamino-2,4,6-trinitrobenzene Inferred by Reactive MD Simulations and Unsupervised Learning

When high-energy-density materials are subjected to thermal or mechanical insults at extreme conditions (shock loading), a coupled response between the thermo-mechanical and chemical behaviors is systematically induced. Herein we develop a reaction model for the fast chemistry of 1,3,5-triamino-2,4,6-trinitrobenzene (TATB) at the mesoscopic scale, where the chemical behavior is determined by underlying microscopic reactive simulations. The slow carbon cluster formation is not discussed in the present work. All-atom reactive molecular dynamics (MD) simulations are performed with the ReaxFF potential, and a reduced-order chemical kinetics model for TATB is fitted to isothermal and adiabatic simulations of single crystal chemical decomposition. Unsupervised machine learning techniques based on non-negative matrix factorization are applied to MD trajectories to model the decomposition kinetics of TATB in terms of a four-component model. The associated heats of reaction are fit to the temperature evolution from adiabatic decomposition trajectories. Using a chemical species analysis, we show that non-negative matrix factorization captures the main chemical decomposition steps of TATB and provides an accurate estimation of their evolution with temperature. The final analytical formulation, coupled to a diffusion term, is incorporated into a continuum formalism, and simulation results are compared one-to-one against MD simulations of 1D reaction propagation along different crystallographic directions and with different initial temperatures. A good agreement is found for both the temporal and spatial evolution of the temperature field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine Learning to Identify Geologic Factors Associated with Production in Geothermal Fields: A Case-Study Using 3D Geologic Data from Brady Geothermal Field and NMFk

In this paper, we present an analysis using unsupervised machine learning (ML) to identify the key geologic factors that contribute to the geothermal production in Brady geothermal field. Brady is a hydrothermal system in northwestern Nevada that supports both electricity production and direct use of hydrothermal fluids. Transmissive fuid-fow pathways are relatively rare in the subsurface, but are critical components of hydrothermal systems like Brady and many other types of fuid-fow systems in fractured rock. Here, we analyze geologic data with ML methods to unravel the local geologic controls on these pathways. The ML method, non-negative matrix factorization with k-means clustering (NMFk), is applied to a library of 14 3D geologic characteristics hypothesized to control hydrothermal circulation in the Brady geothermal field. Our results indicate that macro-scale faults and a local step-over in the fault system preferentially occur along production wells when compared to injection wells and non-productive wells. We infer that these are the key geologic characteristics that control the through-going hydrothermal transmission pathways at Brady. Our results demonstrate: (1) the specific geologic controls on the Brady hydrothermal system and (2) the efficacy of pairing ML techniques with 3D geologic characterization to enhance the understanding of subsurface processes. This submission includes the published journal article detailing this work, the published 3D geologic map of the Brady Geothermal Area used as a basis to develop structural and geological variables that are hypothesized to control or effect permeability or connectivity, 3D well data, along which geologic data were sampled for PCA analyses, and associated metadata file. This work was done using the GeoThermalCloud framework, which is part of SmartTensors (both are linked below).

15 GEOTHERMAL ENERGY↗

Direct high-altitude observations of 2-methyltetrols in the gas- and particle phase in air masses from Amazonia

We present direct observations of 2-methyltetrol (C 5 H 12 O 4 ) in the gas- and particle phase from the deployment of a Filter Inlet for Gases and Aerosols coupled to a Time-of-Flight Chemical Ionization Mass Spectrometer (FIGAERO-CIMS) during the Southern Hemisphere High Altitude Experiment on Particle Nucleation and Growth (SALTENA), which took place between December 2017 and June 2018 at the high-altitude Global Atmosphere Watch station Chacaltaya (CHC) located at 5240 m a s l in the Bolivian Andes. 2-Methyltetrol signals were dominant in a factor resulting from Positive Matrix Factorization (PMF) identified as influenced by Amazon emissions. We combine these observations with investigations of isoprene oxidation chemistry and uptake in an isolated deep convective cloud in the Amazon using a photochemical box model with coupled cloud microphysics and show that, likely, 2-methyltetrol is taken up by hydrometeors or formed in situ in the convective cloud, and then transported in the particle phase in the cold environment of the Amazon outflow and to the station, where it partially evaporates.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An Ensemble Approach to Computationally Efficient Radiological Anomaly Detection and Isotope Identification

Radiological source search is a challenging task involving detection and identification of weak sources in a constantly changing radiological background. As of now, many radiological source detection algorithms have been proposed; however, their computational complexity, and hence reliance on power intensive processing units inhibit low-power applications of radiological source search systems. In this work, we introduce the anomaly filter (AF) algorithm; a computationally light, yet effective time-series source detection algorithm based on exponential weighted moving average (EWMA) and Poisson deviance statistics. Then, we demonstrate that the proposed algorithm can be used in ensemble with other more computationally intensive source detection and identification algorithms to achieve both increased detection performance and reduced power consumption. The proposed AF algorithm and the ensemble algorithms were thoroughly benchmarked against several existing source detection and identification algorithms. The results show that the AF algorithm outperforms existing conventional source detection algorithms, and the ensemble approach improves the overall performance of existing source detection and isotope identification algorithms. Furthermore, the AF algorithm and the Non-negative Matrix Factorization approach based source identification (NMF-ID) algorithm were combined and implemented on a singleboard microcontroller and the power consumption was measured. This ensemble algorithm reduced the power consumption of the NMF-ID algorithm almost by a factor of 100, while improving the detection performance of the overall system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Correlations Between Panoramic Imagery and Gamma-Ray Background in an Urban Area

When searching for radiological sources in an urban area, a vehicle-borne detector system will often measure complex, varying backgrounds primarily from natural gamma-ray sources. Much work has been focused on developing spectral algorithms that retain sensitivity and minimize the false-positive rate even in the presence of such spectral and temporal variability. However, information about the environment surrounding the detector system might also provide useful clues about the expected background, which if incorporated into an algorithm, could improve performance. Recent work has focused on extensive measuring and modeling of urban areas with the goal of understanding how these complex backgrounds arise. This work presents an analysis of panoramic video images and gamma-ray background data collected in Oakland, California, by the radiological multisensor analysis platform (RadMAP) vehicle. Features were extracted from the panoramic images by semantically labeling the images and then convolving the labeled regions with the detector response. A linear model was used to relate the image-derived features to gamma-ray spectral features obtained using nonnegative matrix factorization (NMF) under different regularizations. Here we find some gamma-ray background features correlate strongly with image-derived features that measure the response-adjusted solid angle subtended by sky and buildings, and we discuss the implications for the development of future, contextually aware detection algorithms.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗