Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Computation of rarefied hypersonic flows

Numerical techniques for the simulation of hypersonic flows of rarefied gases are examined in an analytical review. The direct-simulation Monte Carlo (DSMC) method developed to interpret measurement data obtained by the Space Shuttle in the SUMS project is described; the fundamental limitations of the DMSC approach are discussed; and modifications to improve the physical plausibility of DSMC predictions are proposed. Particular attention is given to the use of a double-peaked molecular distribution function for the internal flow in the SUMS probe, a downstream vacuum-reservoir condition as a simplifying assumption, and an explicit forward-time centered-space differencing scheme for the discretization of the SUMS problem. Typical simulation results are presented in extensive graphs and briefly characterized.

Cheng, Sin-I↗

A physical basis for cosmological correlators from cuts

Significant progress has been made in our understanding of the analytic structure of FRW wavefunction coefficients, facilitated by the development of efficient algorithms to derive the differential equations they satisfy. Moreover, recent findings indicate that the twisted cohomology of the associated hyperplane arrangement defining FRW integrals overestimates the number of integrals required to define differential equations for the wave-function coefficient. We demonstrate that the associated dual cohomology is automatically organized in a way that is ideal for understanding and exploiting the cut/residue structure of FRW integrals. Utilizing this understanding, we develop a systematic approach to organize compatible sequential residues, which dictates the physical subspace of FRW integrals for any n -site, ℓ-loop graph. In particular, the physical subspace of tree-level FRW wavefunction coefficients is populated by differential forms associated to cuts/residues that factorize the integrand of the wavefunction coefficient into only flat space amplitudes. After demonstrating the validity of our construction using intersection theory, we develop simple graphical rules for cut tubings that enumerate the space of physical cuts and, consequently, differential forms without any calculation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗

Strategies for concurrent processing of complex algorithms in data driven architectures

Research directed at developing a graph theoretical model for describing data and control flow associated with the execution of large grained algorithms in a special distributed computer environment is presented. This model is identified by the acronym ATAMM which represents Algorithms To Architecture Mapping Model. The purpose of such a model is to provide a basis for establishing rules for relating an algorithm to its execution in a multiprocessor environment. Specifications derived from the model lead directly to the description of a data flow architecture which is a consequence of the inherent behavior of the data and control flow described by the model. The purpose of the ATAMM based architecture is to provide an analytical basis for performance evaluation. The ATAMM model and architecture specifications are demonstrated on a prototype system for concept validation.

Stoughton, John W.↗

Shock associated noise of inverted-profile coannular jets. II - Condition for minimum noise. III - Shock structure and noise characteristic

The generation of noise by the shock-turbulence interaction with shock cells in an inverted-profile coannular jet with nozzle exit velocity aligned with the jet axis is investigated analytically, interpreting the optical measurements of Tanna et al. (1985). The noise-intensity minimum at slightly supersonic primary-flow velocities is related to the weakness of the primary-stream shock-cell structure and the lack of such a pattern in the outer fan stream. The discrepancies between this finding and those of Dosanjh et al. (1977 and 1978) are attributed to nozzle design, the definition of minimum noise, and different interpretative approaches. A first-order shock-cell model is then developed to derive formulas for the peak frequencies and the scaling of noise intensity. The results of computations using these formulas are presented in graphs and found to be in good agreement with the experimental data.

Tam, C. K. W.↗

Quasi-steady flight to quasi-steady flight transition in a windshear - Trajectory guidance

The control (via the angle of attack) of the vertical flight of an aircraft taking off at maximum power in a horizontal wind shear with downdraft is investigated analytically. Optimal trajectories to recover the initial path inclination or to recover quasi-steady flight (the relative values of the velocity, path inclination, and angle of attack) are derived using a Chebyshev approach and shown to be nearly identical in the shear but divergent after the shear. These results are then applied to construct a trajectory-guidance control comprising a variable-gamma guidance scheme for the shear trajectory, a constant-gamma guidance scheme for the immediate postshear trajectory, and a constant-rate-of-climb guidance scheme for the aftershear trajectory. Numerical results demonstrating the near-optimal performance of the control are presented in tables and graphs.

Miele, A.↗

Strategies for concurrent processing of complex algorithms in data driven architectures

The results of ongoing research directed at developing a graph theoretical model for describing data and control flow associated with the execution of large grained algorithms in a spatial distributed computer environment is presented. This model is identified by the acronym ATAMM (Algorithm/Architecture Mapping Model). The purpose of such a model is to provide a basis for establishing rules for relating an algorithm to its execution in a multiprocessor environment. Specifications derived from the model lead directly to the description of a data flow architecture which is a consequence of the inherent behavior of the data and control flow described by the model. The purpose of the ATAMM based architecture is to optimize computational concurrency in the multiprocessor environment and to provide an analytical basis for performance evaluation. The ATAMM model and architecture specifications are demonstrated on a prototype system for concept validation.

Stoughton, John W.↗

An integrated approach to optimizing concentration shock wave electrodialysis using 2D multicell simulation and response surface models

Shock wave electrodialysis (SWED) is a highly promising technique for energy-efficient ion separation in the context of a circular economy. This paper presents a approach way of modeling and improving SWED using a two-dimensional multicell model combined with the COMSOL program and response surface methodology. The model integrates the Nernst-Planck equation, Darcy's law, and first-order electroosmosis to examine the local concentration, flux of ionic species, distribution of current, and velocity of flow in SWED cells under various operating conditions. We first illustrate the clear depiction of concentration, velocity, and electric potential distribution through contours which aids in identifying optimal operating conditions and designing scalable SWED systems. The results emphasize the significance of surface charge density and voltage in influencing the features of shock waves for obtaining effective ion separation while optimizing energy consumption and improving current efficiency by controlling the retention time of feed flow. Here, this study defines two crucial characteristics of shock waves, namely the length of the flat depletion zone of a fully developed shock wave (shock wave height) and the distance of shock wave propagation (shock wave length). These properties significantly impact separation performance, as determined by the simulation results. Additionally, the response surface methodology is incorporated with the COMSOL models to develop predictive models and graph responses, enabling a more comprehensive understanding of the interactions between parameters and performance indicators, such as removal ratio, energy consumption, and water recovery. Finally, this work suggests design tactics for expanding SWED processes and outlines potential areas for further research. This research provides valuable insights into the prospective applications, design optimization, and scalability of SWED in the field of electrokinetic separation technologies for green chemistry and a circular economy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Open Source Software Prevalence Ingest Tool

The OSSP Ingest Tool accepts user-input organizational information, ingests IT/OT asset lists in Excel format, and ingests the associated CycloneDX SBOM's. It then performs analytics demonstrating the ability to answer the follow research questions: o RQ1. Ability to identify all OSS services running on, and all OSS components present within, an OT device o RQ1a: Ability to differentiate multiple versions of the same OSS component within each OT device. o RQ1b: Ability to differentiate running from not-running OSS components. o RQ1c: Ability to differentiate based on the originator of the component, because a supplier may have modified it after retrieval from the upstream software source. o RQ2. Ability to correlate the identity of a single OSS component across multiple OT devices, mitigating common name variations such as differences in capitalization, '-' vs '_', and so on. o RQ3. Ability to perform subset analysis of OSS components across multiple OT devices o RQ3a: Ability to perform subset analysis across OSS libraries, generating density & distribution graphs to identify commonly-used libraries and outliers. o RQ3b: Ability to perform subset analysis of a single OSS library, generating density & distribution by CI sector, by device type, by device make/model, and/or by firmware version. o RQ3c: Ability to perform subset analysis by grouping OSS libraries according to programming language, then overlay with RQ4b. o RQ3d: Ability to perform subset analysis by OSS upstream source, providing insight into degree of modifications performed by suppliers. o RQ4. Ability to identify dependencies (transitive and direct) of each differentiated OSS library within each OT device, and enable RQ1,2,3 iteratively for dependencies. o RQ1. Ability to identify all OSS services running on, and all OSS components present within, an OT device o RQ1a: Ability to differentiate multiple versions of the same OSS component within each OT device. o RQ1b: Ability Page

Kapadia, Shayna [Lawrence Livermore National Labor↗

Determining the Most Influencing Medical Conditions in MEDPRAT’s SIN Directed Graph

INTRODUCTION: The Susceptibility Inference Network (SIN) is a network of medical conditions, part of the Medical Extensible Probabilistic Risk Assessment Tool (MEDPRAT) developed by NASA to assess human health and medical risk to space exploration missions. The SIN is subject matter expert informed and acts as a prototype that provides relationships and dependencies between events modeled by MEDPRAT. Each vertex in the SIN has a weight which evaluates the severity of having the condition regardless of the progression from or to that condition. In this presentation, we consider two statistics to measure that stand alone risk: Quality Time Lost (QTL) and Loss of Crew Life (LOCL). Our goal is to identify the medical conditions that contribute the most to crew members QTL and LOCL risks due to progression of conditions in the network. We investigate how different computation parameters result in different condition rankings and address the choice of parameters that allows appropriate interventions to ensure space mission success. METHODS: The Katz score, one of many centrality measures created for ranking purposes in network analysis, takes into account all possible walks through the network, penalizing each additional step in a walk by a factor α called the Katz parameter. The literature does not provide specific values for the choice of α. We derive an analytical relationship between α and the maximum path length which has influence on the Katz score and ranking. Based on the probability of progression of each condition in the SIN, we identify that maximum path length of interest and calculate α that is then used in the Katz formula to rank the conditions in the SIN. RESULTS AND CONCLUSION: The effective probabilities of the SIN matrix generally fall below 10−6, which is below the level of the least influencing condition in the set. This corresponds to the probability of at most six consecutive progressions of a condition. Consequently, we calculate the Katz Parameter α and get 0.32. We rank the medical conditions and find that Acute Radiation Symptom is the condition the most prone to contribute to quality time loss due to progression.

risk analysis↗

Affect of Brush Seals on Wave Rotor Performance Assessed

The NASA Lewis Research Center's experimental and theoretical research shows that wave rotor topping can significantly enhance gas turbine engine performance levels. Engine-specific fuel consumption and specific power are potentially enhanced by 15 and 20 percent, respectively, in small (e.g., 400 to 700 hp) and intermediate (e.g., 3000 to 5000 hp) turboshaft engines. Furthermore, there is potential for a 3- to 6-percent specific fuel consumption enhancement in large (e.g., 80,000 to 100,000 lbf) turbofan engines. This wave-rotor-enhanced engine performance is accomplished within current material-limited temperature constraints. The completed first phase of experimental testing involved a three-port wave rotor cycle in which medium total pressure inlet air was divided into two outlet streams, one of higher total pressure and one of lower total pressure. The experiment successfully provided the data needed to characterize viscous, partial admission, and leakage loss mechanisms. Statistical analysis indicated that wave rotor product efficiency decreases linearly with the rotor to end-wall gap, the square of the friction factor, and the square of the passage of nondimensional opening time. Brush seals were installed to further minimize rotor passage-to-cavity leakage. The graph shows the effect of brush seals on wave rotor product efficiency. For the second-phase experiment, which involves a four-port wave rotor cycle in which heat is added to the Brayton cycle in an external burner, a one-dimensional design/analysis code is used in conjunction with a wave rotor performance optimization scheme and a two-dimensional Navier-Stokes code. The purpose of the four-port experiment is to demonstrate and validate the numerically predicted four-port pressure ratio versus temperature ratio at pressures and temperatures lower than those that would be encountered in a future wave rotor/demonstrator engine test. Lewis and the Allison Engine Company are collaborating to investigate wave rotor integration in an existing turboshaft engine. Recent theoretical efforts include simulating wave rotor dynamics (e.g., startup and load-change transient analysis), modifying the one-dimensional wave rotor code to simulate combustion internal to the wave rotor, and developing an analytical wave rotor design/analysis tool based on macroscopic balances for parametric wave rotor/engine analysis.

Source record↗

Rocket-Plume Spectroscopy Simulation for Hydrocarbon-Fueled Rocket Engines

The UV-Vis spectroscopic system for plume diagnostics monitors rocket engine health by using several analytical tools developed at Stennis Space Center (SSC), including the rocket plume spectroscopy simulation code (RPSSC), to identify and quantify the alloys from the metallic elements observed in engine plumes. Because the hydrocarbon-fueled rocket engine is likely to contain C2, CO, CH, CN, and NO in addition to OH and H2O, the relevant electronic bands of these molecules in the spectral range of 300 to 850 nm in the RPSSC have been included. SSC incorporated several enhancements and modifications to the original line-by-line spectral simulation computer program implemented for plume spectral data analysis and quantification in 1994. These changes made the program applicable to the Space Shuttle Main Engine (SSME) and the Diagnostic Testbed Facility Thruster (DTFT) exhaust plume spectral data. Modifications included updating the molecular and spectral parameters for OH, adding spectral parameter input files optimized for the 10 elements of interest in the spectral range from 320 to 430 nm and linking the output to graphing and analysis packages. Additionally, the ability to handle the non-uniform wavelength interval at which the spectral computations are made was added. This allowed a precise superposition of wavelengths at which the spectral measurements have been made with the wavelengths at which the spectral computations are done by using the line-by-line (LBL) code. To account for hydrocarbon combustion products in the plume, which might interfere with detection and quantification of metallic elements in the spectral region of 300 to 850 nm, the spectroscopic code has been enhanced to include the carbon-based combustion species of C2, CO, and CH. In addition, CN and NO have spectral bands in 300 to 850 nm and, while these molecules are not direct products of hydrocarbon-oxygen combustion systems, they can show up if nitrogen or a nitrogen compound is present as an impurity in the propellants and/or these can form in the boundary layer as a result of interaction of the hot plume with the atmosphere during the ground testing of engines. Ten additional electronic band systems of these five molecules have been included into the code. A comprehensive literature search was conducted to obtain the most accurate values for the molecular and the spectral parameters, including Franck-Cordon factors and electronic transition moments for all ten band systems. For each elemental transition in the RPSSC, six spectral parameters - Doppler broadened line width at half-height, pressure-broadened line width at half-height, electronic multiplicity of the upper state, electronic term energy of the upper state, Einstein transition probability coefficient, and the atomic line center - are required. Input files have been created for ten elements of Ni, Fe, Cr, Co, Cu, Ca, Mn, Al, Ag, and Pd, which retain only relatively moderate to strong transitions in 300 to 430 nm spectral range for each element. The number of transitions in the input files is 68 for Ni; 148 for Fe; 6 for Cr; 87 for Co; 1 for Ca; 3 for Mn; 2 each for Cu, Al, and Ag; and 11 for Pd.

Tejwani, Gopal D.↗

Active causal learning for decoding chemical complexities with targeted interventions

Abstract Predicting and enhancing inherent properties based on molecular structures is paramount to design tasks in medicine, materials science, and environmental management. Most of the current machine learning and deep learning approaches have become standard for predictions, but they face challenges when applied across different datasets due to reliance on correlations between molecular representation and target properties. These approaches typically depend on large datasets to capture the diversity within the chemical space, facilitating a more accurate approximation, interpolation, or extrapolation of the chemical behavior of molecules. In our research, we introduce an active learning approach that discerns underlying cause-effect relationships through strategic sampling with the use of a graph loss function. This method identifies the smallest subset of the dataset capable of encoding the most information representative of a much larger chemical space. The identified causal relations are then leveraged to conduct systematic interventions, optimizing the design task within a chemical space that the models have not encountered previously. While our implementation focused on the QM9 quantum-chemical dataset for a specific design task—finding molecules with a large dipole moment—our active causal learning approach, driven by intelligent sampling and interventions, holds potential for broader applications in molecular, materials design and discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Self-Supervised T-GCN for Detection of Disturbance and Propagation in Power Grid

Urban power systems increasingly rely on dense sensing to monitor grid reliability, yet disturbance labels are scarce and events are rare. We present a self-supervised spatio-temporal method that detects, localizes, and characterizes grid frequency disturbances across urban areas using only unlabeled data. Our approach trains a tiny Temporal Graph Convolutional Network (T-GCN) to forecast per-site frequency residuals (deviation from 60 Hz). The sensor graph is constructed directly from signals using pre-event Pearson correlation with a cross-correlation lag penalty without geocoding. At inference, node-level anomalies are the model's forecast errors; region-level alarms arise from connected components of high-score nodes. We estimate disturbance propagation by computing per-node arrival times (first persistent exceedance), then fit a planar or time-of-arrival model to obtain direction, speed, and an epicenter proxy. With only three real events collected at decisecond resolution across U.S. cities, we evaluate the T-GCN and report time-to-detect, footprint size, and propagation consistency. We further show that short-window embeddings from the T-GCN's hidden states enable few-shot event-vs-background recognition via a simple prototypical classifier. Despite minimal data and no labels, our system yields fast, spatially coherent detection and interpretable propagation maps, offering a practical, lightweight pathway to city-scale grid resilience analytics.

Niu, Haoran [ORNL] (ORCID:0000000155228297)↗

Tribometer for Lubrication Studies in Vacuum

The NASA Lewis Research Center has developed a new way to evaluate the liquid lubricants used in ball bearings in space mechanisms. For this evaluation, a liquid lubricant is exercised in the rolling contact vacuum tribometer shown in the photo. This tribometer, which is essentially a thrust bearing with three balls and flat races, has contact stresses similar to those in a typical preloaded, angular contact ball bearing. The rotating top plate drives the balls in an outward-winding spiral orbit instead of a circular path. Upon contact with the "guide plate," the balls are forced back to their initial smaller orbit radius; they then repeat this spiral orbit thousands of times. The orbit rate of the balls is low enough, 2 to 5 rpm, to allow the system to operate in the boundary lubrication regime that is most stressful to the liquid lubricant. This system can determine the friction coefficient, lubricant lifetime, and species evolved from the liquid lubricant by tribodegradation. The lifetime of the lubricant charge is only few micrograms, which is "used up" by degradation during rolling. The friction increases when the lubricant is exhausted. The species evolved by the degrading lubricant are determined by a quadrupole residual gas analyzer that directly views the rotating elements. The flat races (plates) and 0.5-in.-diameter balls are of a configuration and size that permit easy post-test examination by optical and electron microscopy and the full suite of modern surface and thin-film chemical analytical techniques, including infrared and Raman microspectroscopy and x-ray photoelectron spectroscopy. In addition, the simple sphere-on-a-flat-plate geometry allows an easy analysis of the contact stresses at all parts of the ball orbit and an understanding of the frictional energy losses to the lubricant. The analysis showed that when the ball contacts the guide plate, gross sliding occurs between the ball and rotating upper plate as the ball forced back to a smaller orbit radius. The friction force due to gross sliding is sensed by the piezoelectric force transducer behind the guide plate and furnishes the coefficient of friction for the system. This tribometer has been used to determine the relative lifetimes of Fomblin Z-25, a lubricant often used in space mechanisms, as a function of the material of the plates against which it was run. The balls were 440C steel in all cases; the plate materials were aluminum, chromium (Cr), 440C steel (17 wt % Cr), and 4150 steel (1 wt % Cr). As shown in the bar graph, the lifetime is greatest for the plate material with least chromium, thus implicating chromium as a tribochemically active element attacking Fomblin Z-25.

Pepper, Stephen V.↗

Characterizing Orbital Debris and Spacecrafts Through a Multi-Analytical Approach

Defining the risks present to both crewed and robotic spacecrafts is part of NASA s mission, and is critical to keep these resources out of harms way. Characterizing orbital debris is an essential part of this mission. We present a proof-of-concept study that employs multiple techniques to demonstrate the efficacy of each approach. The targets of this study are IDCSPs (Initial Defense Communications Satellite Program). 35 of these satellites were launched by the US in the mid-1960s and were the first US communications satellites in the GEO regime. They were emplaced in slightly sub-synchronous orbits. These targets were chosen for this proof-of-concept study for the simplicity of their observable exterior surfaces. The satellites are 26-sided polygons (86cm in diameter), initially spin-stabilized and covered on all sides in solar panels. Data presented here include: (a) visible broadband photometry (Johnson B and Cousins R bands) taken with the University of Michigan s 0.6-m aperture Curtis-Schmidt telescope MODEST (for Michigan Orbital DEbris Survey Telescope) in Chile in November, 2011, (b) laboratory broadband photometry (Johnson BV Cousins RI) of solar cells, obtained using the Optical Measurements Center (OMC) at NASA/JSC (see Cowardin et al., this meeting for more details), (c) visible-band spectra taken using the Magellan 6.5m Baade Telescope at Las Campanas Observatory in Chile in March, 2012 (see also Seitzer et al., this meeting), and (d) visible-band laboratory spectra of solar cells using a Field Spectrometer. Color-color plots using broadband photometry (e.g. B-R vs. R-I) demonstrate that different material types fall into distinct areas on the plots (Cowardin, AMOS 2010). Spectra will be binned in wavelength to compare with photometry results and plotted on the same graph for comparison. This allows us to compare lab data with telescopic data, and photometric results with spectroscopic results. In addition, the spectral response of solar cells in the visible wavelength regime varies from relatively flat (modern black solar cells with uniform albedo as a function of wavelength) to older solar cells whose reflectivity is sharply peaked in the blue (similar to the IDCSP solar cells). With a target like IDCSPs, the material type is known a priori. Therefore, this study will also be used to determine whether laboratory spectra of pre-launch (pristine) solar cells differ from the telescopic spectra of IDCSPs that have been exposed to the harsh environment of space for ~45 years to investigate whether space weathering effects are evident.

Lederer, S. M.↗

Architector 2.0: Expanded Capabilities for Metal Complex Engineering

Automated three-dimensional molecular construction from two-dimensional graph representations is critical to high-throughput discovery eIorts. Software capabilities in this area have accelerated research across fields ranging from protein design and drug discovery to transition metal catalyst development. When Architector was first introduced, it uniquely enabled high-throughput, chemically relevant three-dimensional construction of f-element complexes. Since its introduction, Architector has been applied in large-scale computational campaigns, targeted studies in critical mineral extraction, and artificial intelligence-driven discovery eIorts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

CoarsenConf: Equivariant Coarsening with Aggregated Attention for Molecular Conformer Generation

Molecular conformer generation (MCG) is an important task in cheminformatics and drug discovery. The ability to efficiently generate low-energy 3D structures can avoid expensive quantum mechanical simulations, leading to accelerated virtual screenings and enhanced structural exploration. Several generative models have been developed for MCG, but many struggle to consistently produce high-quality conformers for meaningful downstream applications. To address these issues, we introduce CoarsenConf, which coarse-grains molecular graphs based on torsional angles and integrates them into an SE(3)-equivariant hierarchical variational autoencoder. Through equivariant coarse-graining, we aggregate the fine-grained atomic coordinates of subgraphs connected via rotatable bonds, creating a variable-length coarse-grained latent representation. Our model uses a novel aggregated attention mechanism to restore fine-grained coordinates from the coarse-grained latent representation, enabling efficient generation of accurate conformers. Furthermore, we evaluate the chemical and biochemical quality of our generated conformers on multiple downstream applications, including property prediction and large-scale oracle-based protein docking. Overall, CoarsenConf generates more accurate conformer ensembles compared to prior generative models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗