Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “REDUNDANCY”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Significant acceleration of solid-state NMR simulations via three-angle powder averaging

The anisotropic frequency shifts imparted onto the NMR resonance frequency depend on the spherical angular coordinates that describe the orientations of the NMR interaction tensors with respect to the applied magnetic field direction. Experiments performed using magic-angle spinning, however, gain a dependence on a third angle: the rotor phase γ. Traditionally, a carousel average is performed to integrate over γ, which leads to a slow convergence of intensities without contributing to the underlying powder patterns. Herein, we show an order of magnitude acceleration in computation time may be obtained by including the γ-averaging into the main powder average to eliminate redundant calculation of resonance frequencies.

magic-angle spinning↗

Online task-space motion control for positioner-coordinated multi-robot manufacturing systems

Incorporating multiple robotic manipulators into large-scale manufacturing systems enhances production efficiency and expands manufacturing capabilities beyond those of single-robot systems. Workpiece positioners in robotic manufacturing have demonstrated significant benefits for process optimization, but coordination strategies for multi-robot systems with shared positioners have received limited attention. This work presents a task-space coordinated trajectory-tracking control framework for multi-robot manufacturing systems, in which robots coordinate their motions within a shared, dynamic workpiece positioning frame. A workpiece positioner actively adjusts the pose of the manufactured component to enable greater operational concurrency and improve overall production efficiency. The proposed motion-coordination scheme employs a distributed and scalable architecture, supporting coordination across heterogeneous multi-robot systems. Two optimization methodologies are introduced to manage kinematic redundancies and maintain continuous, near-optimal operation throughout the manufacturing process. The first strategy exploits a task-space dimensionality reduction to achieve locally optimal configurations by leveraging symmetry-axis rotations of the tool. The second strategy utilizes the workpiece positioner to drive the coordinated robots toward stable and kinematically favorable configurations. For both optimization strategies, multiple objectives are defined to improve key performance metrics, including manipulability, configuration consistency, proximity to mechanical limits, and motion efficiency. Addressing a key limitation of existing coordination approaches, the framework is designed around online setpoint modification, allowing coordinated robots to respond effectively to in-situ process feedback. The proposed control framework is validated using the Robot Operating System (ROS) middleware on a combination of physical and simulated multi-robot system hardware.

Arbogast, Alex [ORNL] (ORCID:0000000154740723)↗

A comparative study of multimodal data fusion strategies for planetary spectroscopy

Integrating heterogeneous data sources can improve scientific inference when different modalities capture complementary information, but doing so is challenging in high-dimensional, small-sample settings. In spectroscopy for planetary exploration, Laser-Induced Breakdown Spectroscopy (LIBS), Raman Spectroscopy (Raman), Visible Infrared Spectroscopy (VISIR), and Mid-Infrared Spectroscopy (MIR) each examine different aspects of composition and mineralogy, raising fundamental questions about when and how data fusion improves predictive performance. Using a Mars-relevant set of geologic standards with measurements from all four modalities, we present a rigorous systematic evaluation of four data fusion strategies: low-level (data) fusion, mid-level (feature) fusion, high-level (decision) fusion, and residual-boosting (sequential) fusion. We assess performance in predicting oxide composition via nested cross-validation and corrected significance testing to evaluate whether data fusion improves upon single-modality baselines. We show that data fusion does not uniformly improve accuracy, and that observed gains are modest, oxide-dependent, and sensitive to modality and model structure. To move beyond aggregate accuracy metrics, we use model coefficients, permutation importance, and residual gain analysis to examine how the fusion models weight individual modalities and to identify patterns of apparent complementarity or redundancy. Though focused on spectroscopy for planetary exploration, our framework for data fusion evaluation and interpretation extends to other scientific domains with heterogeneous and scarce data and provides a principled approach evaluating data fusion strategies, interpreting modality contributions, and understanding tradeoffs among data fusion strategies.

97 MATHEMATICS AND COMPUTING↗

Gold-Standard Chemical Database 137 (GSCDB137): A Diverse Set of Accurate Energy Differences for Assessing and Developing Density Functionals

We present GSCDB137, a rigorously curated benchmark library of 137 data sets (8377 entries) covering main-group and transition-metal reaction energies and barrier heights, (intra- and intermolecular) noncovalent interactions, dipole moments, polarizabilities, electric-field response energies, and vibrational frequencies. Legacy data from GMTKN55 and MGCDB84 have been updated to today's best reference values; redundant or low-quality points were removed, and many new, property-focused sets were added. Testing 29 popular density functional approximations (DFAs) confirms the expected Jacob's-ladder hierarchy overall but also reveals notable exceptions: functional performance for frequencies and electric-field properties correlates poorly with that for other ground-state energetics. ωB97M-V and ωB97X-V are the most balanced hybrid meta-GGA and hybrid GGA, respectively; B97M-V and revPBE-D4 lead the meta-GGA and GGA classes. Double hybrids lower mean errors by about 30% versus their hybrid analogues but demand careful frozen-core, basis set, and spin contamination treatment. GSCDB137 offers a comprehensive, openly documented platform for rigorous validation of DFA and universal machine learning potentials, and training of the next generation of exchange-correlation functionals.

Liang, Jiashu [University of California, Berkeley,↗

Improved Accuracy in Semi-Experimental Structure Determination by Resolving Problems Associated with Rotation of Principal Inertial Axes of Isotopologues: Structures of 1,3-Oxazole ( c -C 3 H 3 NO)

The rotational spectrum of the normal isotopologue of 1,3-oxazole (c-C 3 H 3 NO) was observed from 43 to 750 GHz. Over 3900 transitions for the ground vibrational state are measured, assigned, and least-squares fit to sextic centrifugally distorted-rotor Hamiltonians. The measured frequencies and resulting spectroscopic constants from this extended spectral range, combined with previous measurements of the nuclear quadrupole coupling constants, will facilitate astronomical searches for oxazole across the majority of the range of modern radiotelescopes. Spectra for a set of 30 oxazole isotopologues, which include multiple isotopic substitutions of each atom, are used to determine the first semi-experimental equilibrium ($r$$^{SE}_{e}$) structure and semi-experimental substitution structure ($r$$^{SE}_{e}$), each using CCSD(T) computed values for the vibration–rotation interaction and electron-mass corrections. The large number of isotopologues, including 21 isotopologues observed for the first time, and the redundant substitutions of each atom provide sufficient spectroscopic information to determine the $r$$^{SE}_{e}$ structure with the expected high level of accuracy and precision (0.0001 or 0.0002 Å in bond distances and 0.013 to 0.025° in bond angles). In the course of this study, we analyzed a known issue for some $r$$^{SE}_{e}$ structure determinations of near-oblate asymmetric tops in which inclusion of individual isotopologues degrades the structure determination. We demonstrate that this problem primarily arises from the difference in the values of the computed vibration–rotation interaction corrections as evaluated at the computed re geometry vs the $r$$^{SE}_{e}$ geometry of the “real” molecule. Our solution to this problem substantially improves the $r$$^{SE}_{e}$ structure of oxazole and likely can be generalized to many other molecules.

Chemical structure↗

Topology-Informed Design Rules for Deconstructable Thermoset Copolymer Networks

Existing models of thermoset deconstruction facilitated by incorporating cleavable comonomers rely on a mean-field reverse gel point paradigm, which predicts network dissolution once cleavable bonds reach a critical stoichiometric threshold, but does not account for where those bonds reside within the network architecture. Using reactive coarse-grained molecular dynamics simulations coupled with graph-theoretic analysis, we extend this stoichiometric picture to show that deconstructability is governed by the curing-imprinted network topology rather than stoichiometry alone. This topological organization is hierarchical: at the local scale, the elastic effectiveness of cross-link junctions determines which cross-links constitute the load-bearing scaffold; at the mesoscale, the cross-linking rate kinetically templates that scaffold into topologically modular communities─densely cross-linked clusters connected by sparse bridging strands that sustain network connectivity. Using betweenness centrality to identify nodes that disproportionately lie on intercommunity shortest paths, we demonstrate that effective deconstruction of the network into macromolecular fragments requires cleavable comonomers to intercept these high-centrality bridging strands. We further find that under uniform, disassortative comonomer incorporation, this topological requirement provides a mechanistic basis for extending the reverse gel point to incorporate network topology. We also show that modularity imposes a fundamental limit on fragment uniformity that persists even when the centrality requirement is met. Finally, we demonstrate that chain stiffness provides a nearly independent lever to suppress mechanically redundant cross-links and raise the glass transition temperature without significantly altering the deconstruction outcome. Together, these findings reframe the thermoset design space around network topology and provide actionable guidelines for engineering thermoset copolymers with predictable deconstructability and targeted thermomechanical performance.

coarse-grained molecular dynamics↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

Developing Scenario‐Based Strategies for Health, Climate, and Environmental Preparedness: The One Health, One Earth Approach

Climate change amplifies many threats to human health. Despite advances in understanding climate change dynamics and impacts, there remains a critical gap in translating scientific knowledge into equitable, and community-driven health interventions. The inaugural One Earth, One Health workshop sought to explore this gap through human-centered design exercises involving interdisciplinary researchers from climate and Earth sciences, engineering, epidemiology, microbiology, and environmental health. Although participants did not co-develop solutions with affected communities, they used stakeholder role-playing to guide ideation and lay groundwork for actionable plans. Through these methods, participants identified community needs and proposed prototype solutions to alleviate health threats exacerbated by global environmental change. Prototypes were organized around infectious diseases, extreme weather, and air quality, as illustrative themes rather than an exhaustive set of risks. Key solutions included strategies for anticipatory systems and early warning (e.g., integrating environmental signals with health data), inclusive communication and infrastructure needs for responding to extreme weather events, and integrated platforms visualizing air quality trends to support tailored, context-aware guidance beyond one-size-fits-all alerts. The workshop highlighted opportunities such as leveraging machine learning, Earth observation, and real-time surveillance to protect communities, but also noted barriers including data quality, technological redundancy, privacy, and governance challenges. Additionally, participants emphasized the need for interdisciplinary teams capable of collaborating across sectors, breaking down silos and addressing gaps in training and education. Overall, the workshop illustrates how process-driven, human-centered approaches can help surface user needs and generate testable prototype concepts, while underscoring the importance of direct community partnership for implementation.

Abadi, Azar M. [University of Alabama, Birmingham,↗

Functional anatomy of zinc finger antiviral protein complexes

Abstract ZAP is an antiviral protein that binds to and depletes viral RNA, which is often distinguished from vertebrate host RNA by its elevated CpG content. Two ZAP cofactors, TRIM25 and KHNYN, have activities that are poorly understood. Here, we show that functional interactions between ZAP, TRIM25 and KHNYN involve multiple domains of each protein, and that the ability of TRIM25 to multimerize via its RING domain augments ZAP activity and specificity. We show that KHNYN is an active nuclease that acts in a partly redundant manner with its homolog N4BP1. The ZAP N-terminal RNA binding domain constitutes a minimal core that is essential for antiviral complex activity, and we present a crystal structure of this domain that reveals contacts with the functionally required KHNYN C-terminal domain. These contacts are remote from the ZAP CpG binding site and would not interfere with RNA binding. Based on our dissection of ZAP, TRIM25 and KHNYN functional anatomy, we could design artificial chimeric antiviral proteins that reconstitute the antiviral function of the intact authentic proteins, but in the absence of protein domains that are otherwise required for activity. Together, these results suggest a model for the RNA recognition and action of ZAP-containing antiviral protein complexes.

Science & Technology - Other Topics↗

Topobexin targets the Topoisomerase II ATPase domain for beta isoform-selective inhibition and anthracycline cardioprotection

Abstract Topoisomerase II alpha and beta (TOP2A and TOP2B) isoenzymes perform essential and non-redundant cellular functions. Anthracyclines induce their potent anti-cancer effects primarily via TOP2A, but at the same time they induce a dose limiting cardiotoxicity through TOP2B. Here we describe the development of theobexclass of TOP2 inhibitors that bind to a previously unidentified druggable pocket in the TOP2 ATPase domain to act as allosteric catalytic inhibitors by locking the ATPase domain conformation with the capability of isoform-selective inhibition. Through rational drug design we have developed topobexin, which interacts with residues that differ between TOP2A and TOP2B to provide inhibition that is both selective for TOP2B and superior to dexrazoxane. Topobexin is a potent protectant against chronic anthracycline cardiotoxicity in an animal model. This demonstration of TOP2 isoform-specific inhibition underscores the broader potential to improve drug specificity and minimize adverse effects in various medical treatments.

Science & Technology - Other Topics↗

Optimal invariant sets for atomistic machine learning

The representation of atomic configurations for machine learning models has led to numerous sets of descriptors. However, many descriptor sets are incomplete and/or functionally dependent. Incomplete sets cannot faithfully represent atomic environments. Yet complete constructions often suffer from a high degree of functional dependence, where some descriptors are functions of others. These redundant descriptors do not improve discrimination between atomic environments. We employ pattern recognition techniques to remove dependent descriptors to produce the smallest possible set that satisfies completeness. We apply this in two ways: First, we refine an existing description, the atomic cluster expansion. Second, we augment an incomplete construction, yielding a new message-passing neural network architecture that can recognize up to 5-body patterns. This architecture shows strong accuracy on state-of-the-art benchmarks while retaining low computational cost. Our results demonstrate the utility of this strategy to optimize descriptor sets across a range of descriptors and application datasets.

97 MATHEMATICS AND COMPUTING↗

Low-overhead transversal fault tolerance for universal quantum computation

Fast, reliable logical operations are essential for realizing useful quantum computers. By redundantly encoding logical qubits into many physical qubits and using syndrome measurements to detect and correct errors, we can achieve low logical error rates. However, for many practical quantum error correction codes such as the surface code, owing to syndrome measurement errors, standard constructions require multiple extraction rounds—of the order of the code distance d—for fault-tolerant computation, particularly considering fault-tolerant state preparation. Here we show that logical operations can be performed fault-tolerantly with only a constant number of extraction rounds for a broad class of quantum error correction codes, including the surface code with magic state inputs and feedforward, to achieve ‘transversal algorithmic fault tolerance’. Through the combination of transversal operations7 and new strategies for correlated decoding, despite only having access to partial syndrome information, we prove that the deviation from the ideal logical measurement distribution can be made exponentially small in the distance, even if the instantaneous quantum state cannot be made close to a logical codeword because of measurement errors. We supplement this proof with circuit-level simulations in a range of relevant settings, demonstrating the fault tolerance and competitive performance of our approach. Furthermore, our work sheds new light on the theory of quantum fault tolerance and has the potential to reduce the space–time cost of practical fault-tolerant quantum computation by over an order of magnitude.

Zhou, Hengyun [QuEra Computing, Boston, MA (United↗

Coassembly and binning of a twenty-year metagenomic time-series from Lake Mendota

Abstract The North Temperate Lakes Long-Term Ecological Research (NTL-LTER) program has been extensively used to improve understanding of how aquatic ecosystems respond to environmental stressors, climate fluctuations, and human activities. Here, we report on the metagenomes of samples collected between 2000 and 2019 from Lake Mendota, a freshwater eutrophic lake within the NTL-LTER site. We utilized the distributed metagenome assembler MetaHipMer to coassemble over 10 terabases (Tbp) of data from 471 individual Illumina-sequenced metagenomes. A total of 95,523,664 contigs were assembled and binned to generate 1,894 non-redundant metagenome-assembled genomes (MAGs) with ≥50% completeness and ≤10% contamination. Phylogenomic analysis revealed that the MAGs were nearly exclusively bacterial, dominated by Pseudomonadota (Proteobacteria, N = 623) and Bacteroidota (N = 321). Nine eukaryotic MAGs were identified by eukCC with six assigned to the phylum Chlorophyta. Additionally, 6,350 high-quality viral sequences were identified by geNomad with the majority classified in the phylum Uroviricota. This expansive coassembled metagenomic dataset provides an unprecedented foundation to advance understanding of microbial communities in freshwater ecosystems and explore temporal ecosystem dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Security constrained optimal power shutoff for wildfire risk mitigation

Abstract Electric grid faults are increasingly the source of ignition for major wildfires. To reduce the likelihood of such ignitions in high risk situations, utilities use preemptive de‐energization of power lines, commonly referred to as Public Safety Power Shutoffs (PSPS). Besides raising challenging trade‐offs between power outages and wildfire safety, PSPS removes redundancy from the network at a time when component faults are likely to happen. This may leave the network particularly vulnerable to unexpected line faults that may occur while the PSPS is in place. Previous works have not explicitly considered the impacts of these outages. To address this gap, the Security Constrained Optimal Power Shutoff problem is proposed which uses post‐contingency security constraints to model the impact of unexpected line faults when planning a PSPS. This model enables, for the first time, the exploration of a wide range of trade‐offs between both wildfire risk and pre‐ and post‐contingency load shedding when designing PSPS plans, providing useful insights for utilities and policy makers considering different approaches to PSPS. The efficacy of the model is demonstrated using the EPRI 39‐bus system as a case study. The results highlight the potential risks of not considering security constraints when planning PSPS and show that incorporating security constraints into the PSPS design process improves the resilience of current PSPS plans.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Federated Access from DOE Labs to Distributed Storage in the EIC Era of Computing

The Electron Ion Collider (EIC) collaboration and future experiment is a unique scientific ecosystem within Nuclear Physics as the experiment starts right off as a crosscollaboration between Brookhaven National Lab (BNL) & Jefferson Lab (JLab). As a result, this muti-lab computing model tries at best to provide services accessible from anywhere by anyone who is part of the collaboration. While the computing model for the EIC is not finalized, it is anticipated that the computational and storage resources will be made accessible to a wide range of collaborators across the world. The use of federated ID seems to be a critical element to the strategy of providing such services, allowing seamless access to each lab site computing resources. However, providing Federated access to a Federated storage is not a trivial matter and has its share of technical challenges. In this contribution, we focus on the steps we took towards the deployment of a distributed object storage system that integrates with Amazon S3 and Federated ID. We will first cover for and explain the first stage storage solutions provided to the EIC during the detector design phase. Our initial test deployment consisted of Lustre storage using MinIO, hence providing an S3 interface. High Availability load balancers were added later to provide the initial scalability it lacked. Performance of that system will be shown. While this embryonic solution worked well, it had many limitations. Looking ahead, the Ceph object storage is considered a top-of-the-line solution in the storage community - since the Ceph Object Gateway is compatible with the Amazon S3 API out of the box, our next phase will use a native S3 storage. Our Ceph deployment will consist of erasure coded storage nodes to maximize storage potential along with multiple Ceph Object Gateways for redundant access. We will compare performance of our next stage implementations. Finally, we will present how to leverage OpenID Connect with the Ceph Object Gateway’s to enable Federated ID access. We hope this contribution will serve the community needs as we move forward with cross-lab collaborations and the need for Federated ID access to distributed compute facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Prototype-Wise Sensitivity Analysis of Urban Building Energy Simulation Surrogate Modeling Accuracy

Urban Building Energy Modeling (UBEM) is an important reference for urban energy-related policymaking. Because of the significant impact of urban microclimates on the energy simulation, UBEM requires simulations of many microclimate-prototype pairs. Surrogate modeling is commonly used to reduce the cost of simulation computations. In UBEM surrogate modeling, it is important to determine the percentage of microclimates related to a prototype used for generating surrogate model training data. This study analyzes the prototype-wise variations and sensitivities of surrogate model estimation accuracy to the microclimate sampling ratios. The results of the study can help determine the number of simulations used for generating surrogate modeling data, avoid redundant simulations, and reduce the computational cost for UBEM surrogate modeling and its time.

Pan, Xiyu↗

Scaling and merging time-resolved pink-beam diffraction with variational inference

Time-resolved x-ray crystallography (TR-X) at synchrotrons and free electron lasers is a promising technique for recording dynamics of molecules at atomic resolution. While experimental methods for TR-X have proliferated and matured, data analysis is often difficult. Extracting small, time-dependent changes in signal is frequently a bottleneck for practitioners. Recent work demonstrated this challenge can be addressed when merging redundant observations by a statistical technique known as variational inference (VI). However, the variational approach to time-resolved data analysis requires identification of successful hyperparameters in order to optimally extract signal. In this case study, we present a successful application of VI to time-resolved changes in an enzyme, DJ-1, upon mixing with a substrate molecule, methylglyoxal. We present a strategy to extract high signal-to-noise changes in electron density from these data. Furthermore, we conduct an ablation study, in which we systematically remove one hyperparameter at a time to demonstrate the impact of each hyperparameter choice on the success of our model. We expect this case study will serve as a practical example for how others may deploy VI in order to analyze their time-resolved diffraction data.

47 OTHER INSTRUMENTATION↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗