Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

81 records · Page 5

Convergent Concordant Mode Approach for Molecular Vibrations: CMA-2

The concordant mode approach (CMA) is a promising new scheme for dramatically increasing the system size and level of theory achievable in quantum chemical computations of molecular vibrational frequencies. Here, we achieve advances in the CMA hierarchy by computations targeting CCSD(T)/cc-pVTZ (coupled cluster singles and doubles with perturbative triples using a correlation-consistent polarized-valence triple-ζ basis set) benchmarks within the G2 molecular test set, executing a statistical analysis for 1501 frequencies from 111 compounds and then separately solving the refractory case of pyridine. First, MP2/cc-pVTZ (second-order Møller–Plesset perturbation theory with the same basis set) proves to be an excellent and preferred choice for generating the underlying (Level B) normal modes of the CMA scheme. Utilizing this Level B within the CMA-0A method reproduces the 1501 benchmark frequencies with a mean absolute error (MAE) of only 0.11 cm –1 and an attendant standard deviation of 0.49 cm –1 . Second, a convergent CMA-2 method is constituted that allows efficient computation of higher level (Level A) frequencies to any reasonable accuracy threshold by using only Hartree–Fock (HF) and MP2 or density functional theory (DFT) data to generate ξ parameters, which select the sparse off-diagonal force field elements for explicit evaluation at Level A. When Level B = MP2/cc-pVTZ, a cutoff of ξ = 0.02 provides an average maximum absolute error per molecule of only 0.17 cm –1 by incurring merely a 33% increase in average cost over CMA-0A. This CMA-2 method also eradicates the 4 problematic CMA-0A outliers of pyridine with even less effort (ξ = 0.04, 22% increase). Finally, the newly developed CMA procedures are shown to be highly successful when applied to 1-(1H-pyrrol-3-yl)ethanol, a new test molecule with diverse types of vibration.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reduced Order Model to Predict Dispersion of Flammable Refrigerant into a Space

As the HVAC&R industry mobilizes to deploy more low-GWP refrigerants, relevant standards are being continually reviewed and updated. Those include the general safety standards ISO 5149 and ASHRAE 15, and the equipment standards IEC and UL. The standards systematically set the allowable maximum amount of refrigerant that should be used in different equipment types and different applications. To do so, they rely on predictions of how a leaked refrigerant mass will disperse into a space. Dispersion characteristics, such as total flammable volume and its residence time, determine the risk associated with the presence of the flammable refrigerant. The standards have included provisions for the use of flammable refrigerants for approximately two decades. They relied on limited analytical analyses and test cases in their development. Dispersion of a refrigerant into a space is complex. Computational fluid dynamics (CFD) are the most accurate in predicting a given problem. However, CFD is computationally expensive and requires specialized expertise and resources and is not suitable for use by standards development working group as prediction tool. This paper presents the development of a reduced order model (ROM) that predicts the key dispersion characteristics relevant to the dispersion of a leaked refrigerant into a space for any combination of input variables. The inputs are the refrigerant release height, the total released refrigerant mass and its release flow rate, the refrigerant molecular weight, the ventilation flow rate, the floor area and height of the space, recirculation air flow rate, and the tightness of the space. The outputs are histograms of volume fraction of the room in prescribed concentration bins and the total mass of the refrigerant in each bin normalized by the total refrigerant charge at 13 prescribed simulation time stamps between 1 and 900 seconds. The ROM is constructed from a set of CFD simulations with carefully chosen combinations of input parameters. The selection if done using a multidimensional sparse grid which is a generalization of the classical tensor approach but offers additional flexibility and thus can be more carefully tuned towards a specific model. The tuning is done to improve the accuracy, measured in the difference between the output values of the ROM and the CFD model, while minimizing the computational cost, measured in number of CFD simulations which is orders of magnitude more expensive than the processing the training data.

Edwards, Dean↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

Implicit Formulation of Muscle Dynamics in OpenSim

Astronauts lose bone and muscle mass during spaceflight. Exercise countermeasure is the primary method for counteracting bone and muscle mass loss in space. New spacecraft exercise device concepts are currently being developed for the NASAs new crew exploration vehicle. The NASA Digital Astronaut Project (DAP) uses computational modeling to help determine if the new exercise devices will be effective as countermeasures. The NASA Digital Astronaut Project is developing the ability to utilize predictive simulation to provide insight into the change in kinematics and kinetics with a change in device and gravitational environment (1-g versus 0-g). For example, in space exercise the subject's body weight is applied in addition to the loads prescribed for musculoskeletal maintenance. How and where these loads are applied obviously directly impacts bone and tissue loads. Additionally, due to space vehicle structural requirements, exercise devices are often placed on vibration isolation systems. This changes the apparent impedance or stiffness of the device as seen by the user. Data collection under these conditions is often impractical and limited. Predictive modeling provides a means to have a virtual subject to test hypotheses. Predictive simulation provides a virtual subject for which we are able to perform studies such as sensitivity to device loading and vibration isolation without the need for laboratory kinematic or kinetic test data.Direct Collocation optimization provides an efficient means to perform task based optimization and predictive modeling. It is relatively straight forward to structure a physical exercise task in a Direct Collocation mathematical formulation: perform a motion such that you start at an initial pose, achieve a given amount of deflection i.e a squat, return to the initial pose, and minimize muscle activation cost. Direct Collocation is advantageous in that it does not require numerical integration to evaluate the objective function. Instead, the system dynamics are transformed to discrete time and the optimizer is constrained such that the solution is not considered to be a valid unless the dynamic equations are satisfied at all time points. The simulation and optimization are effectively done simultaneously. Due to the implicit integration, time steps can be more coarse than in a differential equation solver. In a gait scenario this means that that the model constraints and cost function are evaluated at 100 nodes in the gait cycle versus 10,000 integration steps in a variable-step forward dynamic simulation. Furthermore, no time is wasted on accurate simulations of movements that are far from the optimum. Constrained optimization algorithms require a Jacobian matrix that contains the partial derivatives of each of the dynamic constraints with respect to of each of the state and control variables at all time points. This is a large but sparse matrix. An implicit dynamics formulation requires computation of the dynamic residuals f as a function of the states x and their derivatives, and controls u:f(x, dxdt, u) 0If the dynamics of musculoskeletal system are formulated implicitly, the Jacobian elements are often available analytically, eliminating the need for numerical differentiation; this is obviously computationally advantageous. Additionally, implicit formulation of musculoskeletal dynamics do not suffer from singularities from low mass bodies, zero muscle activation, or other stiff system or

physical exercise↗

Annotation of DOM metabolomes with an ultrahigh resolution mass spectrometry molecular formula library

Current approaches to analyzing metabolomic data often rely on matching MS/MS fragmentation data to sparse libraries or databases. This approach results in limited identification of features, often with less than 10% of the dataset being annotated. A complementary approach is to assign molecular formula to features based on accurate mass measurements, but the platforms commonly used for metabolomics do not have the needed accuracy or resolving power to do this robustly, particularly for larger molecules. Using our newly modified analysis tool, CoreMS, we generated a library of molecular formula from pooled samples analyzed with LC-21T FT-ICR MS. This library successfully annotated approximately 53.2% of features identified from the exometabolome of marine diatom Phaeodactylum tricornutum – a nearly ten-fold increase over the 5.9% annotation rate achieved using a conventional MS/MS library matching approach. Using this FT-ICR MS library approach, we were able to differentiate differences in the exometabolome of P. tricornutum in iron replete and iron limited conditions, with 668 metabolites being differentially expressed (p < 0.05, 2 x intensity difference) under these conditions. The traditional MS/MS fragmentation-based annotation approach only annotated 61 of these metabolites, while our novel pipeline annotated 450 metabolites and revealed 12 metabolites that were significantly more abundant under low iron conditions. Our results demonstrate the utility of ultrahigh resolution mass spectrometry for generating more comprehensive and confident molecular annotations.

21T-FTICR-MS, CoreMS↗

Testing to Evaluate Processes Expected to Occur during MSR Salt Spill Accidents

Obtaining a license for a new nuclear reactor requires the identification and assessment of the potential consequences of specified accident scenarios, which are achieved using accident progression models. Those models need to be parameterized and validated using experimental data, but existing experimental data addressing processes relevant to molten salt reactor accidents are sparse. Specifically, experimental data that quantify the sensitives of important processes to the initial conditions of the spill, the ambient environment, and the containment features are needed to parameterize individual process models. Integrated experiments that simulate accident scenarios are also needed to provide data for model validation, but these experiments will require the use of proven methods to quantify the processes under evaluation. The overarching objectives of this work are to develop the methods for simulating the targeted processes, to determine the effectiveness of the methods in producing the data required for model development, to generate data that can be used to parameterize individual process models, and to provide key insights into the behavior of spilled molten salt that should be considered in models. Experimental methods were designed to quantify aspects of individual processes expected to occur during or after a molten salt spill accident that will affect the fate of spilled molten salt and the radionuclides within. These processes include 1) molten salt spreading and heat transfer, 2) molten salt flowing and freezing in tubing, 3) stainless steel corrosion kinetics in molten salt, and 4) molten salt splashing and aerosol generation. The initial tests described in this report were conducted using eutectic FLiNaK to demonstrate the test methods, the data that are generated, and the analyses of the data to derive values needed for modeling. The primary variables that were tested include initial salt temperature and the presence of volatile surrogate fission products (e.g., cesium and iodine). The developed methods are shown to be effective in quantifying the desired processes and can be applied to study more complex salt compositions of interest to molten salt reactor developers, a wide range of environmental conditions of interest to modelers, and additional variables relevant to salt spill accidents. The developed methods and insights gained from laboratory tests can also be incorporated in future large-scale integrated tests used to simulate molten salt spill accidents.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Proteome-wide analysis of protein stability in Escherichia coli under acid stress

Knowledge of protein acid sensitivity remains sparse and is largely derived from low-throughput, enzyme-specific assays. We used a scalable framework to map acid stability across the Escherichia coli proteome to assess the acid stability of 1,675 unique proteins, estimating pH 50 values for over 90% of them. The parameter pH50 was defined as the pH value at which only 50% of the initial protein remains in solution following acid treatment. Proteome-wide pH 50 values ranged from 2.28 to 6.33 (median 5.11). Approximately 9% of detected proteins remained stable across all tested pH conditions. Our results align with published data and the assay of citrate synthase (GltA) performed here. Protein acid stability differed significantly by subcellular localization: periplasmic proteins were relatively more abundant in the acid-stable group, cytoplasmic proteins were abundant at pH 50 values 4.5–5.5, and inner membrane proteins at higher pH 50 between 5.5 and 6.0. Outer membrane proteins were too few to draw strong conclusions regarding enrichment within specific pH 50 groups. Notably, the periplasmic binding protein of the molybdate ABC transporter (ModA), was enriched after incubation at low pH. Estimated pH 50 values showed no correlation with protein isoelectric point and molecular weight. Together, this work provides the first proteome-wide map of protein acid stability and establishes a general framework for studying different chemical stressors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Novel Application of Machine Learning Techniques for Rapid Source Apportionment of Aerosol Mass Spectrometer Datasets

In this work, we apply machine learning approaches sparse multinomial logistic regression to classify aerosol mass spectrometer (AMS) unit mass resolution (UMR) data followed by an ensemble regression technique for source apportionment of organic aerosols (OA). The classifier was trained on 60 well characterized laboratory and positive matrix factorization (PMF) deconvolved reference spectra to identify eight OA types. These include four laboratory-derived secondary organic aerosol (SOA) spectra, which include isoprene photooxidation SOA, isoprene epoxydiols (IEPOX) SOA, a monoterpene SOA type that includes a-pinene and ß-pinene SOA, and aromatic SOA from oxidation of naphthalene and m-xylene precursors, as well as PMF deconvolved spectra for three primary organic aerosol (POA) types, namely, hydrocarbon-like organic aerosol (HOA), biomass burning organic aerosol (BBOA), and cooking OA (COA), and a more oxidized oxygenated OA type (MO-OOA). A 5-fold cross-validation strategy, repeated 10 times, was used to assess the classifier’s performance. The classifier had high classification accuracy for COA, aromatic SOA, and isoprene SOA spectra but incorrectly classified ~9% by number of MO-OOA spectra as BBOA, 12% of BBOA spectra as HOA (and vice versa), and 18% of IEPOX-SOA spectra as aromatic SOA. Next, an ensemble regression model was trained on an artificially generated dataset consisting of mixtures of different OA types to assess its ability to predict fractional mass abundances from classification probabilities of various OA species obtained from the multinomial logistic regression classifier trained on the reference spectra. Ultimately, the proposed approach was applied for source apportionment of aircraft-based AMS measurements of OA UMR spectra during the HI-SCALE field campaign. On two representative days (May 6th and 18th, 2016), the algorithm determined that ~50-60% of OA by mass was MO-OOA, which represented a highly aged organic aerosol mixture from different sources. On both days, BBOA was determined to contribute less than 10% to OA by mass. However, on May 18th, the aromatic SOA fraction was higher compared to that on May 6th. The proposed approach is capable of rapidly analyzing AMS data in real time, making it suitable for applications where rapid source apportionment of AMS OA spectra is desirable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗