Engineering PapersSearch

SEARCH · Engineering Papers

Results for “embedding model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Modeling low cycle fatigue (LCF) of additively manufactured Hastelloy X using An accelerated crystal plasticity fatigue damage model

This paper presents a microstructure-based model for low cycle fatigue (LCF) behavior and life of Nickel-based alloy Hastelloy X manufactured using laser-powder bed fusion (L-PBF) additive manufacturing (AM). AM Hastelloy X, a solution-strengthened alloy, is tested at elevated temperature under fully reversed LCF conditions at different strain levels. A generalized plane strain finite element model is generated from electron backscatter diffraction (EBSD) characterization. The constitutive behavior of the material under fatigue is modeled using crystal plasticity and calibrated with both monotonic tensile and cyclic stress–strain data. The fatigue micro-crack initiation and propagation in the microstructure is modeled using a modified Chaboche fatigue damage model. An embedded boundary condition with a homogenous medium is used to apply the cyclic deformation and prevent numerically introduced over-constraints during fatigue simulation. A ‘cycle-jump’ method is used to accelerate the fatigue simulation and reduce the computational cost. The simulation results are compared to LCF experiments, showing satisfactory matches in cyclic stress behavior and number of cycles to macro-crack initiation for all applied strain ranges. In addition, the model illustrates the potential for quantifying microscale fatigue life impacting factors such as microstructure and surface roughness, which is needed to accurately quantify the reliability of AM components in service.

36 MATERIALS SCIENCE

Assessment of uranium nitride interatomic potentials

Uranium mononitride (UN) is a promising nuclear fuel due to its high fissile density, high thermal conductivity, and suitability for reprocessing. In this study, two uranium nitride interatomic potentials are assessed: Tseplyaev and Starikov's angular-dependent potential and Kocevski et al.'s embedded atom model potential. Predictions of the thermophysical and elastic properties of UN, UN 2 , and α- and β-U 2 N 3 computed using both potentials are assessed and compared to available experimental data. Notably, the Tseplyaev potential performs better with the energetic aspects of UN, e.g., specific heat capacity and point defect formation energies, whereas the Kocevski potential performs better with the structural aspects of UN, e.g., thermal expansion as well as with the elastic properties. The reasons why the Kocevski potential underestimates the UN specific heat are explained by examining the UN phonon properties modeled using both potentials. The Kocevski potential shows better identification of the mechanical stability ranges of UN, UN 2 , and α- and β-U 2 N 3 , reasonably predicting the melting point of UN and predicting stable structures for UN 2 and α- and β-U 2 N 3 . On the other hand, the Tseplyaev potential predicts a premature phase change of both UN and UN 2 and cannot stabilize α- nor β-U 2 N 3 . However, the Kocevski potential cannot predict a stable α-U phase and is thus not suitable for the calculation of formation energies for non-stoichiometric point defects.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES

Thermal Modeling and Limitations for Power Electronics Embedded in Medium-Voltage Cables

As next-generation energy technologies gain traction and power demand increases, the existing electrical infrastructure faces significant stress, prompting innovative solutions to enhance the grid's capacity and lifespan. This work explores the possibility of embedding medium-voltage (MV) power electronics directly inline with the cable, and the resulting thermal challenges. Since the majority of power distribution cables installed in the U.S. are passively cooled, the work focuses primarily on passive cooling, with an emphasis on the limitations of axial heat spreading within the cable. To date, literature on axial spreading of high incident heat loads on cables and cable environments is limited, typically reporting cases with <10 W of incident heat load. This work will explore the considerations, limits, and tradeoffs of cable-embedded heat loads significantly larger than the cable losses. Both external and internal effects are modeled analytically in nondimensional terms via a Biot number analysis, allowing fundamental limits and tradeoffs to be derived. The work culminates in the design and experimental validation of a cable-embedded thermal system capable of passively dissipating 300 W of heat from a coaxial SiC mosfet switch module over a length of 20 cm, thus validating the possibility of MV cable-embedded power electronics from a thermal standpoint.

24 POWER TRANSMISSION AND DISTRIBUTION

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the representational qualities of larger models in smaller models. We observe a geometric phenomenon which we term embedding condensation, where token embeddings collapse into a narrow cone-like subspace in some language models. Through systematic analyses across multiple Transformer families, we show that small models such as GPT2 and Qwen3-0.6B exhibit severe condensation, whereas larger models such as GPT2-x1 and Qwen3-32B are more resistant to this phenomenon. Additional observations show that embedding condensation is not reliably mitigated by knowledge distillation from larger models. To fight against it, we formulate a dispersion loss that explicitly encourages embedding dispersion during training. Experiments demonstrate that it mitigates condensation, recovers dispersion patterns seen in larger models, and yields performance gains across 10 benchmarks. We believe this work offers a principled path toward improving smaller Transformers without additional parameters.

Xiao, Xi [ORNL] (ORCID:0009000009316982)

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab

Active learning of a crystal plasticity flow rule from discrete dislocation dynamics simulations

Continuum-scale material deformation models, such as crystal plasticity (CP), can significantly enhance their predictive accuracy by incorporating input from lower-scale (i.e. mesoscale) models. The procedure to generate and extract the relevant information is however typically complex and ad hoc, involving decision and intervention by domain experts, leading to long development times. In this study, we develop a principled approach for calibration of continuum-scale models using lower scale information by representing a CP flow rule as a Gaussian process model. This representation allows for efficient parameter space exploration, guided by the uncertainty embedded in the model through a process known as Bayesian optimization (BO). We demonstrate a semi-autonomous BO loop which instantiates discrete dislocation dynamics simulations whose initial conditions are automatically chosen to optimize the uncertainty of a model CP flow rule. Our self-guided computational pipeline efficiently generated a dataset and corresponding model whose error, uncertainty, and physical feature sensitivities were validated with comparison to an independent dataset four times larger, demonstrating a valuable and efficient active learning implementation readily transferable to similar material systems.

36 MATERIALS SCIENCE

Modeling CO 2 flow through faulted/fractured reservoirs using tEDFM in corner-point grids

The interest in underground CO 2 storage has increased significantly over the last decade because of the rising concern about global warming due to the growing levels of greenhouse gases in the atmosphere. Considering that CO 2 accounts for 80% of these greenhouse gases, carbon capture, utilization, and storage (CCUS) is regarded as one of the most direct approaches to achieving the net zero carbon target. Although CO 2 storage in deep saline aquifers and depleted gas reservoirs has been studied extensively, most studies use commercial simulators that model faults/fractures by simply modifying the transmissibility in the direction perpendicular to the fault surfaces. Here, this work shows that this simplistic approach ignores the accelerated flow in the directions parallel to the fault plane, leading to significantly higher leakage along the fault surface. To accurately model the flow of CO 2 in faulted reservoirs, we present the first transient embedded discrete fracture model for corner-point grids (tEDFM-CPG). By comparing the results of the tEDFM-CPG to high-resolution reference solutions, we show that this approach is accurate and efficient at predicting CO 2 flow in faulted/fractured reservoirs. Finally, this work presents the use of mixed reality (MR) to efficiently observe CO 2 gas migration in the interior of these corner-point grid systems.

25 ENERGY STORAGE

OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model

With the sharp increasing volume of user data, Deep Learning Recommendation Model (DLRM) becomes an indispensable infrastructure in large technology companies. However, large-scale DLRM on the multi-GPU platform is still inefficient due to unbalanced workload partitioning and intensive inter-GPU communication. To this end, we propose OPER, an OPtimality guided Embedding table placement for large-scale Recommendation model training and inference. OPER explores the potential of mitigating remote memory access latency in DLRM through fine-grained embedding table placement. Specifically, OPER proposes a theoretical modeling that builds up the relationship between EMT placement and the embedding communication latency in both training and inference. OPER proves the NP hardness of finding the optimal embedding table placement and proposes a heuristic algorithm that yields near optimal placement. OPER implements a SHMEM-based embedding table training system and a unified embedding index mapping to support fine-grained embedding table sharding and placement. Comprehensive experiments reveal that OPER achieves on average 3.4× and 5.1× speedup on training and inference respectively over state-of-the-art DLRM frameworks.

Wang, Zheng

Modeling Co2 Flow Through Faulted/Fractured Reservoirs Using Tedfm in Corner-Point Grids

Interest in underground CO2 storage has increased significantly over the last decade, driven by growing concern about global warming and rising levels of greenhouse gases in the atmosphere. Given that CO2 accounts for 80% of these greenhouse gases, carbon capture, utilization, and storage (CCUS) is considered one of the most direct approaches to achieving the net-zero carbon target. Although CO2 storage in deep saline aquifers and depleted gas reservoirs has been studied extensively, most studies use commercial simulators that model faults/fractures by simply modifying the transmissibility in the direction perpendicular to the fault surfaces. This work shows that this simplistic approach ignores the accelerated flow in the directions parallel to the fault plane, leading to significantly higher leakage along the fault surface. To accurately model CO2 flow in faulted reservoirs, we present the first transient embedded discrete-fracture model for corner-point grids (tEDFM-CPG). By comparing the tEDFM-CPG results with high-resolution reference solutions, we show that this approach is accurate and efficient at predicting CO2 flow in faulted/fractured reservoirs. In conclusion, this work presents the use of mixed reality (MR) to efficiently observe CO2 gas migration in the interior of these corner-point grid systems.

02 PETROLEUM

Fusion Model for Metagenomics

This work highlights the use of an embeddings approach that can encode multiple features and create efficient contextualization of profiled metagenomes derived from microbiome samples using computer vision models and image representations of the abundance profiles. The model's embeddings can be used to cluster existing samples based on multiple conditions and interpretations, and new embeddings can be quickly created for new samples and fitted to existing clusters to characterize them. This has practical applications for unknown, unlabeled microbiome samples. The model's embeddings can be used to cluster existing samples based on multiple conditions and interpretations, and new embeddings can be quickly created for new samples and fitted to existing clusters to characterize them. This has practical applications for unknown, unlabeled microbiome samples.

Valdes, CamiloA [Lawrence Livermore National Labor

Impacts of Control, Penetration, and Distribution of Embedded Storage Network in Bulk Power System

The current shift in generation mix from fossil fuel plants towards variable and intermittent renewable energy sources is poised to create a future grid with reduced physical inertia and mismatch between generation and demand. Embedded storage, which is a concept of a coordinated network of storage units sited at the interface between the transmission and distribution system, is proposed as a mechanism to provide a buffer between generation and demand. This paper proposes an automated framework to model and integrate embedded storage in large-scale power systems with industry-grade grid-following (GFL) and grid-forming (GFM) control technologies. More importantly, the developed framework is used to explore the impacts of embedded storage control, penetration, location, and capacity in providing fast frequency response to the grid under contingency events such as generator trips and faults. The framework and study are conducted using the transient-stability simulation tool PSS/E and a realistic model of the Puerto Rico grid as a chosen test system. The simulation results show that GFL and GFM embedded storage, distributed throughout the system, with sufficient penetration and capacity, can effectively improve primary frequency response of the system under the studied contingency events.

Battery Energy Storage, embedded storage, grid-for

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan

Modeling supercritical CO2 injection induced rupture of a minor fault embedded in a poroelastic layered reservoir-caprock system

CO2 injection for geologic carbon sequestration involves hydromechanical processes that lead to changes in fluid pressure and stresses that can activate existing faults. This paper presents a new method and workflow of modeling fault activation considering more complex three-dimensional geometry of natural faults using the TOUGH-FLAC multiphase fluid flow and geomechanical simulator. In this method and workflow, FLAC3D mechanical interfaces and TOUGH3 finite volume elements are discretized using computer aided design and gridding software along with a tailored mesh translation routine. The method and workflow are demonstrated with a model of a curved minor fault embedded in a poro-elastic layered reservoir-caprock system. The model is used for a comprehensive sensitivity analysis of fault responses to fault length, injection mass rate, injection schedule, well-fault distance, and well locations versus fault location. Four metrics (CO2 plume, shear state of fault, pressure and stress path at fault monitoring points) are selected to assess CO2 migration, pressure change, and the reactivation of faults. The results reveal that CO2 can bypass around the tip of the minor impermeable fault, building up pressure and poro-elastic stress on both sides that tends to impede fault rupture. Our study shows the benefit of carefully designing the injection to achieve the targeted final storage volume, starting at a relatively low rate for considerable time, and then ramping up the injection rate to the full rate of injection. The initial low injection has two distinct benefits: (1) it allows for the formation of an extensive CO2 plume with a much higher mobility through a low viscosity that will result in a lower pressure for a given injection rate, and (2) it allows for gradual build-up of horizontal poro-elastic stress within the reservoir that will tend to impede activation of steeply dipping faults. The injection scenario starting at a low injection rate, denoted here as conservative injection, can significantly reduce the risk of fault activation as high fluid mobility and reservoir strengthening poro-elastic stress has been established long before reaching the peak injection rates. Moreover, simultaneous injection in two injection wells on both sides of fault can provide further reservoir strengthening through poro-elastic stress buildup acting on a fault under normal faulting stress regime. The findings presented in the paper can provide practical and effective guidance on long-term, safe, and reliable geological CO2 storage.

Cao, Meng

Transfer Learning Meets Embedded Correlated Wavefunction Theory for Chemically Accurate Molecular Simulations: Application to Calcium Carbonate Ion Pairing

Achieving chemical accuracy for molecular simulations remains a central challenge in computational chemistry. Here, we present an embedded correlated wavefunction transfer learning (ECW-TL) framework for accurately simulating molecular dynamics in the condensed phase. ECW-TL incorporates high-level electron exchange and correlation effects in ECW theory while preserving the training and computational efficiency of machine-learned interatomic potentials. We demonstrate the framework on Ca 2+ –CO 3 2– ion pairing in aqueous solution, a key process underlying CO 2 mineralization in seawater. As proof of principle, we first show that fine-tuning a DFT-revPBE-D3(BJ) baseline model with embedded-DFT-SCAN data reproduces the DFT-SCAN free-energy surface within 1 kcal/mol across all solvation states. Extending the framework to embedded MP2 and localized natural-orbital CCSD(T) further refines the free-energy profile, revealing the crucial role of exact electron exchange and correlation in determining ion-pair stability and structure. The computed ion-pair association free energy is in quantitative agreement with experimental measurements, further validating the accuracy of the ECW-TL framework. ECW-TL thus provides a general, data-efficient route for transferring CW accuracy to efficient simulations of complex aqueous and interfacial chemical processes.

cluster chemistry

MTL_TX: A Multi-Task Transformer Model for Improved Radiation Time-Series Estimation

Controlling radiation doses at potential radioactive facilities is critical to ensuring the safety of both personnel and the public. At the Thomas Jefferson National Accelerator Facility (JLab), multiple sensors are deployed around the three experimental halls to monitor key parameters, including single-beam current, energy levels, current leakage, and radiation values during accelerator operations. In this study, we developed a Multi-task Transformer model, MTL_TX, to accurately estimate radiation doses at sensor locations based on historical data, with the aim of enhancing safety in accelerator facilities and surrounding public areas. To improve estimation accuracy, we integrated two innovative components into the proposed model: hierarchical feature embedding (HFE) and multi-level decomposition attention (MDA). Additionally, the multi-task learning (MTL) framework effectively leverages correlations among multiple sensors, enabling individual estimations for each sensor. MTL_TX achieved outstanding results on data collected in 2018, with an MSE of 0.1464, an RMSE of 0.2353, and an R 2 score of 0.8584. Furthermore, when trained on 2018 data, MTL_TX exhibited excellent generalization capability to unseen datasets from 2016 to 2019, achieving an MSE of 0.1407, an RMSE of 0.2263, and an R 2 score of 0.8831. These results demonstrate a significant improvement over existing state-of-the-art models.

Transformer

Interpretable and flexible non-intrusive reduced-order models using reproducing kernel Hilbert spaces

This paper develops an interpretable, non-intrusive reduced-order modeling technique using regularized kernel interpolation. Existing non-intrusive approaches approximate the dynamics of a reduced-order model (ROM) by solving a data-driven least-squares regression problem for low-dimensional matrix operators. Our approach instead leverages regularized kernel interpolation, which yields an optimal approximation of the ROM dynamics from a user-defined reproducing kernel Hilbert space. We show that our kernel-based approach can produce interpretable ROMs whose structure mirrors full-order model structure by embedding judiciously chosen feature maps into the kernel. The approach is flexible and allows a combination of informed structure through feature maps and closure terms via more general nonlinear terms in the kernel. We also derive a computable a posteriori error bound that combines standard error estimates for intrusive projection-based ROMs and kernel interpolants. In conclusion, the approach is demonstrated in several numerical experiments that include comparisons to operator inference using both proper orthogonal decomposition and quadratic manifold dimension reduction.

Data-driven model reduction