Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “similarity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

The Power of Many: An Ensemble Approach to Spectral Similarity

Quantifying the similarity between two mass spectra─a known reference mass spectrum and an unidentified sample mass spectrum─is at the heart of compound identification workflows in gas chromatography–mass spectrometry (GC-MS). The reference spectrum most like the sample is assigned as its identification (provided some quantitative similarity threshold is met, e.g., 80%) and thus accurately measuring similarity is essential. Significant research has gone toward developing metrics for this purpose, each of which has attempted to improve upon existing methods by incorporating GC-MS-specific information (e.g., peak ratios or retention times) or adopting various statistical and algorithmic frameworks. While this active development has led to a plethora of similarity metrics with demonstrated value across different contexts, the unfortunate consequence has been confusion surrounding which metric should be used as a global standard. No such metric is currently accepted as the standard method because different metrics have demonstrated optimal performance in different contexts. In this work, we propose an ensemble approach to spectral similarity scoring that combines the collective information from across existing similarity metrics to form an improved, globally representative similarity metric as a step toward establishing a global standard method. In conclusion, the resulting ensemble metrics are evaluated on over 88,000 spectra of varying complexity and demonstrate improved abilities to accurately rank the correct reference spectrum as the top-matching candidate for a sample relative to the rankings generated by individual similarity scores.

Carbohydrates↗

Neural architecture search via similarity adaptive guidance

Evolutionary neural network architecture search (ENAS) has attracted the attention of many experts due to its global optimization capabilities to automatically search for convolutional neural network architectures based on the target task. The current search space for ENAS is not to design a fully structured network, but to search for smaller cell architectures to reduce search costs. However, blind search strategies do not effectively utilize the potential experience of the population. In order to utilize the potential experience learned by the current population to guide the evolutionary search of the population, we propose a similarity guided neural network architecture search algorithm based on cell architecture, which utilizes the similarity between pairwise architectures in the population as empirical knowledge learned by the population. Our proposed algorithm provides a novel method for calculating architecture similarity, which calculates architecture similarity separately from the cell and macro-structure. Then we decouple the connections and operations in the cell and calculate connection and operation similarity separately. In addition, we propose adaptive similarity selection and binary tournament selection strategies to enhance the algorithm’s global and local search capabilities and effectively explore the search space. Finally, we design an improved single-point crossover operator to enhance the local search ability of the evolutionary operator. The experimental results show that SAGNAS is a competitive algorithm that achieves 97.44% and 81.60% in CIFAR10 and CIFAR100 with only 1.9 GPU-days spent.

97 MATHEMATICS AND COMPUTING↗

miss-SNF: a multimodal patient similarity network integration approach to handle completely missing data sources

Abstract Motivation Precision medicine leverages patient-specific multimodal data to improve prevention, diagnosis, prognosis, and treatment of diseases. Advancing precision medicine requires the non-trivial integration of complex, heterogeneous, and potentially high-dimensional data sources, such as multi-omics and clinical data. In the literature, several approaches have been proposed to manage missing data, but are usually limited to the recovery of subsets of features for a subset of patients. A largely overlooked problem is the integration of multiple sources of data when one or more of them are completely missing for a subset of patients, a relatively common condition in clinical practice. Results We propose miss-Similarity Network Fusion (miss-SNF), a novel general-purpose data integration approach designed to manage completely missing data in the context of patient similarity networks. miss-SNF integrates incomplete unimodal patient similarity networks by leveraging a non-linear message-passing strategy borrowed from the SNF algorithm. miss-SNF is able to recover missing patient similarities and is “task agnostic”, in the sense that can integrate partial data for both unsupervised and supervised prediction tasks. Experimental analyses on nine cancer datasets from The Cancer Genome Atlas (TCGA) demonstrate that miss-SNF achieves state-of-the-art results in recovering similarities and in identifying patients subgroups enriched in clinically relevant variables and having differential survival. Moreover, amputation experiments show that miss-SNF supervised prediction of cancer clinical outcomes and Alzheimer’s disease diagnosis with completely missing data achieves results comparable to those obtained when all the data are available. Availability and implementation miss-SNF code, implemented in R, is available at https://github.com/AnacletoLAB/missSNF.

Biochemistry & Molecular Biology↗

On the Departure from Monin–Obukhov Surface Similarity and Transition to the Convective Mixed Layer

Large-eddy simulations are used to evaluate mean profile similarity in the convective boundary layer (CBL). Particular care is taken regarding the grid sensitivity of the profiles and the mitigation of inertial oscillations in the simulation spin-up. The nondimensional gradients Φ for wind speed and air temperature generally align with Monin–Obukhov similarity across cases but have a steeper slope than predicted within each profile. The same trend has been noted in several other recent studies. The Businger-Dyer relations are modified here with an exponential cutoff term to account for the decay in Φ to first-order approximation, yielding improved similarity from approximately 0.05z i to above 0.3z i , where z i is the CBL depth. The necessity for the exponential correction is attributed to an extended transition from surface scaling to zero gradient in the mixed layer, where the departure from Monin–Obukhov similarity may be negligible at the surface but becomes substantial well below the conventional surface layer height of 0.1 z i .

54 ENVIRONMENTAL SCIENCES↗

Optimization techniques in self-similar compressible flow

We investigate the one-dimensional (1D) inviscid compressible flow equations for an ideal gas through the lens of optimization techniques. It is the case that, to our knowledge, optimization analysis applied to the so-called “linear velocity” solutions of the Euler compressible flow equations has not been previously conducted. Through both gradient-based and variational techniques, new variants of well-studied flow scenarios, i.e., self-similar, 1D, linear velocity solution class to idealized inviscid compressible flow equations, are determined, as encoded in both the kinematic and thermodynamic properties of this self-similar solution class. With the kinematics of the said solutions being driven by a self-similar “scale radius” and the thermodynamics being driven separately through the appearance of an arbitrary function, a myriad of new solution classes is possible. Acting as a guide to more realistic physical circumstances as well as discovery, it is the hope that the presented cases serve as the framework for future investigations into the intersection of self-similarity and optimization techniques. Fields of study that may find this work to be of interest include aerodynamic design, flow control, inertial confinement fusion, physics-informed neural networks, and other related areas of interest.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Self-similar Reynolds-averaged mechanical–scalar turbulence models for Rayleigh–Taylor mixing induced by power-law accelerations in the small Atwood number limit

Analytical self-similar solutions to two-, three-, and four-equation Reynolds-averaged mechanical–scalar turbulence models describing turbulent Rayleigh–Taylor mixing driven by a temporal power-law acceleration are derived in the small Atwood number (Boussinesq) limit. The solutions generalize those previously derived for constant acceleration Rayleigh–Taylor mixing for models based on the turbulent kinetic energy K and its dissipation rate ε, together with the scalar variance S and its dissipation rate χ [O. Schilling, “Self-similar Reynolds-averaged mechanical–scalar turbulence models for Rayleigh–Taylor, Richtmyer–Meshkov, and Kelvin–Helmholtz instability-induced mixing in the small Atwood number limit,” Phys. Fluids 33, 085129 (2021)]. The turbulent fields are expressed in terms of the model coefficients and power-law exponent, with their temporal power-law scalings obtained by requiring that the self-similar equations are explicitly time-independent. Mixing layer growth parameters and other physical observables are obtained explicitly as functions of the model coefficients and parameterized by the exponent of the power-law acceleration. Values for physical observables in the constant acceleration case are used to calibrate the two-, three-, and four-equation models, such that the self-similar solutions are consistent with experimental and numerical simulation data corresponding to a canonical (i.e., constant acceleration) Rayleigh–Taylor turbulent flow. The calibrated four-equation model is then used to numerically reconstruct the mean and turbulent fields, and turbulent equation budgets across the mixing layer for several values of the power-law exponent. Finally, the reference solutions derived here can be used to understand the model predictions for strongly accelerated or decelerated Rayleigh–Taylor mixing in the large Reynolds number limit.

42 ENGINEERING↗

Towards Automatically Matching Security Advisories to CPEs: String Similarity-based Vendor Matching

When a vulnerability is reported by the National Vulnerability Database (NVD), affected products are listed in the structured Common Platform Enumeration (CPE) format. Unfortunately, if the vulnerability is in a software library (e.g., Log4j), it will not include CPEs for each product containing that library. In these cases, security operators need to manually read the vendor's or third-party security advisories to see if their product is affected. However, these advisories do not report affected products in a structured format, which prevents automated processing, This paper makes the first effort towards automatically constructing structured CPEs for the vulnerable products in a non-NVD security advisory from the unstructured data in the advisory. Since this is a very challenging problem, this paper specifically focuses on the initial but key step of matching the un-structured vendor names in security advisories to the structured vendor representations in the standard CPE format. We explore the feasibility of using string similarity to solve the problem. The basic idea is to compare a vendor name from the non-NVD advisory with each vendor in the official CPE dictionary. The CPE vendor with the highest similarity score to the advisory's vendor will be considered as the match. We first conduct an experimental, comparative study of multiple mainstream string similarity metrics for this matching problem. To improve the performance, we then design a new string similarity metric that is adapted from an existing metric by weighing different tokens in the advisory's vendor name differently.

McClanahan, Kylie↗

Using Hydrodynamic Similarity as a Verification Method for Impact Cratering Simulations in the FLAG Hydrocode

Hydrodynamic codes (hydrocodes) are common tools for modeling hypervelocity impacts to provide insight into the physical phenomenon. Hydrocodes can simulate impacts from micrometer to kilometer spatial scales and reach impact velocities difficult to achieve in experimental settings. However, numerical models are approximations, and demonstrating that a numerical method is capable of providing physical results for these models is essential. In this work, we employ a hydrocode verification technique that leverages hydrodynamic similarity, a mathematical property of the conservation equations of fluid mechanics that form the basis for hydrocode models. Using the FLAG hydrocode, we simulate aluminum (Al) and basalt projectiles and targets at spatial scales spanning 7 orders of magnitude (hundreds of micrometers to kilometers). These materials were chosen because Al-6061 is a common material in spacecraft and satellites and basalt is a useful approximation of rocky astronomical bodies. Our results show that hydrodynamic similarity holds for each material model used and across spatial scales. We show that under certain conditions hydrodynamic similarity can apply in the presence of gravity and that similarity does not hold in the presence of strength models. We conclude that the FLAG hydrocode preserves important mathematical properties of fluid dynamics in hypervelocity impacts of Al-6061 and basalt.

79 ASTRONOMY AND ASTROPHYSICS↗

Similarity for downscaled kinetic simulations of electrostatic plasmas: Reconciling the large system size with small Debye length

A simple similarity has been proposed for kinetic (e.g., particle-in-cell) simulations of plasma transport that can effectively address the long-standing challenge of reconciling the tiny Debye length with the vast system size. This applies to both transport in unmagnetized plasma and parallel transport in magnetized plasmas, where the characteristics length scales are given by the Debye length, collisional mean free paths, and the system or gradient lengths. The controlled scaled variables are the configuration space, x/L, and an artificial Coulomb Logarithm, L ln Λ, for collisions, while the scaled time, t/L, and electric field, LE, are automatic outcomes. The similarity properties are examined, demonstrating that the macroscopic transport physics is preserved through a similarity transformation while keeping the microscopic physics at its original scale of Debye length. To showcase the utility of this approach, two examples of 1D plasma transport problems were simulated using the VPIC code: the plasma thermal quench in tokamaks [Li et al., Nuclear Fusion 63, 066030 (2023)] and the plasma sheath in the high-recycling regime [Li et al., Physics of Plasmas 30, 063505 (2023)].

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Abbreviated Report for 25-FS-011: Switch and Stitch Similar Subgraph Synthesizer

Government institutions utilize software from a wide array of development sources, including those written by large software companies, government contractors, and open-source repositories. Avoiding installation of malicious software components is an important national security endeavor. Automated analysis of previously unseen software is an active research area, and there is much research in the design of systems that compare new software artifacts to a large set of previously seen software records organized by various behaviors they contain. One attractive approach is to turn each compiled software binary into a graph representation and apply a graph similarity model that scores pairs of binaries by their relative similarity. When a pair is deemed similar, it is useful to know why, in the sense of providing explanations to security analysts regarding which portions of the software they should look into further.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Anomaly Detection and Approximate Similarity Searches of Transients in Real-time Data Streams

Abstract We present Lightcurve Anomaly Identification and Similarity Search ( LAISS ), an automated pipeline to detect anomalous astrophysical transients in real-time data streams. We deploy our anomaly detection model on the nightly Zwicky Transient Facility (ZTF) Alert Stream via the ANTARES broker, identifying a manageable ∼1–5 candidates per night for expert vetting and coordinating follow-up observations. Our method leverages statistical light-curve and contextual host galaxy features within a random forest classifier, tagging transients of rare classes ( spectroscopic anomalies), of uncommon host galaxy environments ( contextual anomalies), and of peculiar or interaction-powered phenomena ( behavioral anomalies). Moreover, we demonstrate the power of a low-latency (∼ms) approximate similarity search method to find transient analogs with similar light-curve evolution and host galaxy environments. We use analogs for data-driven discovery, characterization, (re)classification, and imputation in retrospective and real-time searches. To date, we have identified ∼50 previously known and previously missed rare transients from real-time and retrospective searches, including but not limited to superluminous supernovae (SLSNe), tidal disruption events, SNe IIn, SNe IIb, SNe I-CSM, SNe Ia-91bg-like, SNe Ib, SNe Ic, SNe Ic-BL, and M31 novae. Lastly, we report the discovery of 325 total transients, all observed between 2018 and 2021 and absent from public catalogs (∼1% of all ZTF Astronomical Transient reports to the Transient Name Server through 2021). These methods enable a systematic approach to finding the “needle in the haystack” in large-volume data streams. Because of its integration with the ANTARES broker, LAISS is built to detect exciting transients in Rubin data.

79 ASTRONOMY AND ASTROPHYSICS↗

Reassessing the Origins and Contemporary Relevance of ck Acceptability Parameters: Evolving Perspectives on Similarity

“Sensitivity and Uncertainty Analyses Applied to Criticality Safety Validation,” introduces sensitivity and uncertainty methods to address challenges in defining and extending areas of applicability for criticality safety validation. These areas are traditionally defined by the bounds or limits on key parameters, but establishing valid ranges and managing complex parameter variations remain challenging. NUREG/CR-6655 introduces ck and other integral indices, as well as concepts such as the completeness of benchmark coverage, to better quantify system similarities. The work proposed herein seeks to evaluate these foundational concepts to ensure that the bounds remain effective in guiding the assessment of similarity and applicability in modern applications. The concept of completeness, along with other parameters envisioned within the framework, serves as an example of the foundational ideas that have been established, though their effectiveness in practice may not be fully understood. Advancements in scripting tools, coupled with the speed and efficiency of modern computing and statistical models, now allow for faster and more thorough assessments than previously possible. These advancements also enable the identification of trends within the data, which could provide additional insight into system behavior and further broaden the scope of previously performed benchmarks. By leveraging these capabilities, we will revisit and expand the scope of these foundational methods to determine whether the necessary elements for robust similarity evaluation are already embedded, partially realized, or remain untapped.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Compatibility of divertor detachment and ELM suppression in DIII-D high- β p plasmas with ITER-similar shape

Abstract Integration of transient and steady-state divertor heat fluxes control with a high-performance core is necessary for future fusion reactors. In recent DIII-D high- β p experiments, divertor detachment and simultaneous edge localized mode (ELM) suppression are demonstrated while the plasma confinement quality is maintained high in ITER-similar shape. By optimizing the neon injection in high- β p scenario with ITER-similar shape, deep detachment and ELM suppression are achieved with a high-performance core ( β N ∼ 2.8, β p ∼ 2.3) at q 95 ∼ 7.5. Partial divertor detachment and suppression of large ELMs are achieved at q 95 ∼ 6. The stability analyses suggest that with low neon injection, the density pedestal becomes higher and steeper and the T i profile also increases, therefore the increased edge pressure and higher current density destabilize the Peeling-Ballooning mode (PBM), which would lead to a large ELM collapse. With strong neon gas puffing, the significantly reduced pedestal pressure and current density, due to the degraded T e pedestal, lead to the stabilization of PBM and ELMs are suppressed. For both cases, the coupling between the large radius internal transport barrier (ITB) and edge pedestal is the key reason for maintaining high global performance. The formation of large radius ITB compensates for pedestal degradation. Such results could provide an attractive scenario to well control the transient and steady-state heat flux onto the divertor plates while maintaining good plasma performance, which is an important step toward the steady-state operation of future fusion reactors.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

On the transferability of residence time distributions in two 10-km long river sections with similar hydromorphic units

Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.

54 ENVIRONMENTAL SCIENCES↗

Coupled Cluster Theory for Nonadiabatic Dynamics: Nuclear Gradients and Nonadiabatic Couplings in Similarity Constrained Coupled Cluster Theory

Coupled cluster theory is one of the most accurate electronic structure methods for predicting ground and excited state chemistry. However, the presence of numerical artifacts at electronic degeneracies, such as complex energies, has made it difficult to apply the method in nonadiabatic dynamics simulations. While it has already been shown that such numerical artifacts can be fully removed by using similarity constrained coupled cluster (SCC) theory [J. Phys. Chem. Lett. 2017, 8(19), 4801–4807], simulating dynamics requires efficient implementations of gradients and nonadiabatic couplings. Here, we present an implementation of nuclear gradients and nonadiabatic derivative couplings at the similarity constrained coupled cluster singles and doubles (SCCSD) level of theory, thereby making possible nonadiabatic dynamics simulations using a coupled cluster theory that provides a correct description of conical intersections between excited states. We present a few numerical examples that show good agreement with literature values and discuss some limitations of the method.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Nearby Sites Show Similar Upwind Sources and Differing Semivolatile Concentrations in Coastal Aerosol Particles

The Eastern Pacific Cloud Aerosol Precipitation Experiment (EPCAPE) characterized aerosol composition using measurements at two sites within 3 km (Scripps Pier and Mt. Soledad) from 15 February 2023 to 14 February 2024. Comparing the two sites shows the strong influence of upwind sources that results in similar monthly compositions at both sites. The seasonal changes in chemical mass concentrations were largely driven by the upwind source regions, with coastal northwesterly back-trajectories occurring 63−65% of the year and bringing submicrometer mass concentrations that were lower than the EPCAPE average for all trajectories at each site. In contrast, refractory black carbon (rBC) and nonrefractory (NR)-organics and nitrate mass concentrations exceeded EPCAPE average concentrations for back-trajectories from urban areas such as Los Angeles-Long Beach. For hourly measurements, NR-organics and non-sea-salt (NSS)-sulfate mass concentrations at Mt. Soledad were correlated strongly (r = 0.73−0.82) to those measured at Scripps Pier, but NR-nitrate was correlated only moderately (r = 0.63). The explanation for the lower correlation of NR-nitrate is both emissions between the sites and semivolatility, with semivolatility accounting for site-to-site changes in daily averages of +0.01 μg m −3 per percentage site-to-site difference in relative humidity and −0.07 μg m −3 per degree Celsius site-to-site difference in temperature. On average, comparing Scripps Pier to Mt. Soledad, NR-nitrate was higher by 29% because of relative humidity and lower by −26% because of temperature. NR-nitrate and rBC mass concentrations at Scripps Pier for nighttime were 13−15% higher than those for daytime because land breezes brought higher inland concentrations. Concentrations of rBC were 52% higher at Mt. Soledad than those measured at Scripps Pier, accompanied by increases in tracers for brake wear because of traffic on the steep roads within 10 m of that site. The implications are that these nearby sites had comparable monthly concentrations of measured components due to their similar backtrajectories, but hourly and daily concentration differences supported quantification of the meteorological effects from relative humidity and temperature on semivolatile NR-nitrate as well as minor differences from land−sea breezes and local emissions.

Aerosols↗

Galaxy cluster matter profiles - I. Self-similarity, mass calibration, and observable-mass relation validation employing cluster mass posteriors

We present a study of the weak lensing inferred matter profiles ΔΣ(R) of 698 South Pole Telescope (SPT) thermal Sunyaev-Zel’dovich effect (tSZE) selected and MCMF optically confirmed galaxy clusters in the redshift range 0.25 < z < 0.94 that have associated weak gravitational lensing shear profiles from the Dark Energy Survey (DES). Rescaling these profiles to account for the mass dependent size and the redshift dependent density produces average rescaled matter profiles ΔΣ(R/R200c)/(ρcritR200c) with a lower dispersion than the unscaled ΔΣ(R) versions, indicating a significant degree of self-similarity. Galaxy clusters from hydrodynamical simulations also exhibit matter profiles that suggest a high degree of self-similarity, with RMS variation among the average rescaled matter profiles with redshift and mass falling by a factor of approximately six and 23, respectively, compared to the unscaled average matter profiles. We employed this regularity in a new Bayesian method for weak lensing mass calibration that employs the so-called cluster mass posterior P(M200|ζ̂, λ̂, z), which describes the individual cluster masses given their tSZE (ζ̂) and optical (λ̂, z) observables. This method enables simultaneous constraints on richness λ-mass and tSZE detection significance ζ-mass relations using average rescaled cluster matter profiles. We validated the method using realistic mock datasets and present observable-mass relation constraints for the SPT×DES sample, where we constrained the amplitude, mass trend, redshift trend, and intrinsic scatter. Our observable-mass relation results are in agreement with the mass calibration derived from the recent cosmological analysis of the SPT×DES data based on a cluster-by-cluster lensing calibration. Our new mass calibration technique offers a higher efficiency when compared to the single cluster calibration technique. We present new validation tests of the observable-mass relation that indicate the underlying power-law form and scatter are adequate to describe the real cluster sample but that also suggest a redshift variation in the intrinsic scatter of the λ-mass relation may offer a better description. In addition, the average rescaled matter profiles offer high signal-to-noise ratio (S/N) constraints on the shape of real cluster matter profiles, which are in good agreement with available hydrodynamical ΛCDM simulations. This high S/N profile contains information about baryon feedback, the collisional nature of dark matter, and potential deviations from general relativity.Key words: gravitational lensing: weak / galaxies: clusters: general / large-scale structure of Universe

79 ASTRONOMY AND ASTROPHYSICS↗

Spectral Similarity Masks Structural Diversity at Hydrophobic Water Interfaces

The air-water and graphene-water interfaces represent quintessential examples of the liquid-gas and liquid-solid boundaries, respectively. While the sum-frequency generation (SFG) spectra of these interfaces show similarities, a consensus on their signals and interpretations has yet to be reached. Leveraging deep learning, we computed first-principles SFG spectra for both systems, addressing experimental discrepancies. Here, our findings reveal that similarities in SFG signals do not translate into comparable interfacial microscopic properties. Instead, graphene-water and air-water interfaces exhibit fundamental differences in SFG-active thicknesses, hydrogen-bonding networks, and surface dynamics. These distinctions underscore roughness suppression and electronic interactions present at the solid-liquid interface but absent at the gas-liquid interface.

Wang, Yong [Princeton Univ., NJ (United States)] (↗