MOFUN: a Python package for molecular find and replace
MOFUN is an open-source Python package that can find and replace molecular substructures in a larger, potentially periodic, system.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
MOFUN is an open-source Python package that can find and replace molecular substructures in a larger, potentially periodic, system.
High-Energy Resolution Fluorescence Detection X-Ray Fluorescence (HERFD-XRF) imaging and HERFD X-ray Absorption Near Edge Structure (XANES) spectroscopy are used to quantify and characterize trace platinum (Pt) in gold solidi from the Late Roman and Byzantine Empires. Historically, the elemental analysis of coins has been pivotal in distinguishing authentic artifacts from forgeries, elucidating minting practices, and understanding economic shifts. Notably, a new gold source with high platinum content appeared in the fourth century CE, transforming the Roman economy. Traditional methods struggled to detect platinum due to the overwhelming gold matrix. Here, this study demonstrates the effectiveness of HERFD techniques in resolving this challenge. Three gold solidi, minted between 654 and 659 CE, were analyzed alongside reference gold materials with known Pt concentrations. The HERFD-XRF imaging revealed spatial distributions of platinum, highlighting non-uniformities within the coins. Additionally, HERFD-XANES spectroscopy identified the oxidation states and chemical speciation of platinum. Results demonstrate that platinum in the solidi primarily exists as metallic Pt, with some surface oxidation. The findings align with previous measurements but reveal higher Pt concentrations and significant inhomogeneities. This research confirms the reliability of HERFD methods for quantifying trace elements and provides new insights into the raw material sources and minting techniques of ancient gold coins. The non-destructive nature of this approach allows for extensive analyses, offering valuable data for historical, economic, and archaeological studies. This innovative application of HERFD-XRF imaging and XANES in cultural heritage research underscores the potential for detailed material characterization and conservation, enhancing our understanding of ancient economies and trade patterns.
In this work, we present a machine learning approach to calculating electronic specific heat capacities for a variety of benchmark molecular systems. Our models are based on data from density matrix quantum Monte Carlo, which is a stochastic method that can calculate the electronic energy at finite temperature. As these energies typically have noise, numerical derivatives of the energy can be challenging to find reliably. In order to circumvent this problem, we use Gaussian process regression to model the energy and use analytical derivatives to produce the specific heat capacity. From there, we also calculate the entropy by numerical integration. We compare our results to cubic splines and finite differences in a variety of molecules in which Hamiltonians can be diagonalized exactly with full configuration interaction. We finally apply this method to look at larger molecules where exact diagonalization is not possible and make comparisons with more approximate ways to calculate the specific heat capacity and entropy.
Eukaryotic cells form condensates to sense and adapt to their environment [S. F. Banani, H. O. Lee, A. A. Hyman, M. K. Rosen,Nat. Rev. Mol. Cell Biol.18, 285–298 (2017), H. Yoo, C. Triandafillou, D. A. Drummond,J. Biol. Chem.294, 7151–7159 (2019)]. Poly(A)-binding protein (Pab1), a canonical stress granule marker, condenses upon heat shock or starvation, promoting adaptation [J. A. Ribacket al.,Cell168, 1028–1040.e19 (2017)]. The molecular basis of condensation has remained elusive due to a dearth of techniques to probe structure directly in condensates. We apply hydrogen–deuterium exchange/mass spectrometry to investigate the mechanism of Pab1’s condensation. Pab1’s four RNA recognition motifs (RRMs) undergo different levels of partial unfolding upon condensation, and the changes are similar for thermal and pH stresses. Although structural heterogeneity is observed, the ability of MS to describe populations allows us to identify which regions contribute to the condensate’s interaction network. Our data yield a picture of Pab1’s stress-triggered condensation, which we term sequential activation (Fig. 1A), wherein each RRM becomes activated at a temperature where it partially unfolds and associates with other likewise activated RRMs to form the condensate. Subsequent association is dictated more by the underlying free energy surface than specific interactions, an effect we refer to as thermodynamic specificity. Our study represents an advance for elucidating the interactions that drive condensation. Furthermore, our findings demonstrate how condensation can use thermodynamic specificity to perform an acute response to multiple stresses, a potentially general mechanism for stress-responsive proteins.
Abstract Recent advances in scanning tunneling and transmission electron microscopies (STM and STEM) have allowed routine generation of large volumes of imaging data containing information on the structure and functionality of materials. The experimental data sets contain signatures of long-range phenomena such as physical order parameter fields, polarization, and strain gradients in STEM, or standing electronic waves and carrier-mediated exchange interactions in STM, all superimposed onto scanning system distortions and gradual changes of contrast due to drift and/or mis-tilt effects. Correspondingly, while the human eye can readily identify certain patterns in the images such as lattice periodicities, repeating structural elements, or microstructures, their automatic extraction and classification are highly non-trivial and universal pathways to accomplish such analyses are absent. We pose that the most distinctive elements of the patterns observed in STM and (S)TEM images are similarity and (almost-) periodicity, behaviors stemming directly from the parsimony of elementary atomic structures, superimposed on the gradual changes reflective of order parameter distributions. However, the discovery of these elements via global Fourier methods is non-trivial due to variability and lack of ideal discrete translation symmetry. To address this problem, we explore the shift-invariant variational autoencoders (shift-VAEs) that allow disentangling characteristic repeating features in the images, their variations, and shifts that inevitably occur when randomly sampling the image space. Shift-VAEs balance the uncertainty in the position of the object of interest with the uncertainty in shape reconstruction. This approach is illustrated for model 1D data, and further extended to synthetic and experimental STM and STEM 2D data. We further introduce an approach for training shift-VAEs that allows finding the latent variables that comport to known physical behavior. In this specific case, the condition is that the latent variable maps should be smooth on the length scale of the atomic lattice (as expected for physical order parameters), but other conditions can be imposed. The opportunities and limitations of the shift VAE analysis for pattern discovery are elucidated.
ABSTRACT Strongly lensed quadruply imaged quasars (quads) are extraordinary objects. They are very rare in the sky and yet they provide unique information about a wide range of topics, including the expansion history and the composition of the Universe, the distribution of stars and dark matter in galaxies, the host galaxies of quasars, and the stellar initial mass function. Finding them in astronomical images is a classic ‘needle in a haystack’ problem, as they are outnumbered by other (contaminant) sources by many orders of magnitude. To solve this problem, we develop state-of-the-art deep learning methods and train them on realistic simulated quads based on real images of galaxies taken from the Dark Energy Survey, with realistic source and deflector models, including the chromatic effects of microlensing. The performance of the best methods on a mixture of simulated and real objects is excellent, yielding area under the receiver operating curve in the range of 0.86–0.89. Recall is close to 100 per cent down to total magnitude i ∼ 21 indicating high completeness, while precision declines from 85 per cent to 70 per cent in the range i ∼ 17–21. The methods are extremely fast: training on 2 million samples takes 20 h on a GPU machine, and 108 multiband cut-outs can be evaluated per GPU-hour. The speed and performance of the method pave the way to apply it to large samples of astronomical sources, bypassing the need for photometric pre-selection that is likely to be a major cause of incompleteness in current samples of known quads.
ABSTRACT Galaxy clusters enable unique opportunities to study cosmology, dark matter, galaxy evolution, and strongly lensed transients. We here present a new cluster-finding algorithm, CluMPR (Clusters from Masses and Photometric Redshifts), that exploits photometric redshifts (photo-z’s) as well as photometric stellar mass measurements. CluMPR uses a 2D binary search tree to search for overdensities of massive galaxies with similar redshifts on the sky and then probabilistically assigns cluster membership by accounting for photo-z uncertainties. We leverage the deep DESI Legacy Survey grzW1W2 imaging over one-third of the sky to create a catalogue of $\sim 300\, 000$ galaxy cluster candidates out to z = 1, including tabulations of member galaxies and estimates of each cluster’s total stellar mass. Compared to other methods, CluMPR is particularly effective at identifying clusters at the high end of the redshift range considered (z = 0.75–1), with minimal contamination from low-mass groups. These characteristics make it ideal for identifying strongly lensed high-redshift supernovae and quasars that are powerful probes of cosmology, dark matter, and stellar astrophysics. As an example application of this cluster catalogue, we present a catalogue of candidate wide-angle strongly lensed quasars in Appendix C. The nine best candidates identified from this sample include two known lensed quasar systems and a possible changing-look lensed QSO with SDSS spectroscopy. All code and catalogues produced in this work are publicly available (see Data Availability).
We continue examining statistical data assimilation (SDA), an inference methodology, to infer solutions to neutrino flavor evolution, for the first time using real - rather than simulated - data. The model represents neutrinos streaming from the Sun's center and undergoing a Mikheyev-Smirnov-Wolfenstein (MSW) resonance in flavor space, due to the radially-varying electron number density. The model neutrino energies are chosen to correspond to experimental bins in the Sudbury Neutrino Observatory (SNO) and Borexino experiments, which measure electron-flavor survival probability at Earth. In conclusion, the procedure successfully finds consistency between the observed fluxes and the model, if the MSW resonance - that is, flavor evolution due to solar electrons - is included in the dynamical equations representing the model.
Macromolecular crystallography contributes significantly to understanding diseases and, more importantly, how to treat them by providing atomic resolution 3D structures of proteins. This is achieved by collecting X-ray diffraction images of protein crystals from important biological pathways. Spotfinders are used to detect the presence of crystals with usable data, and the spots from such crystals are the primary data used to solve the relevant structures. Having fast and accurate spot finding is essential, but recent advances in synchrotron beamlines used to generate X-ray diffraction images have brought us to the limits of what the best existing spotfinders can do. This bottleneck must be removed so spotfinder software can keep pace with the X-ray beamline hardware improvements and be able to see the weak or diffuse spots required to solve the most challenging problems encountered when working with diffraction images. In this paper, we first present Bragg Spot Detection (BSD), a large benchmark Bragg spot image dataset that contains 304 images with more than 66 000 spots. We then discuss the open source extensible U-Net-based spotfinder Bragg Spot Finder (BSF), with image pre-processing, a U-Net segmentation backbone, and post-processing that includes artifact removal and watershed segmentation. Finally, we perform experiments on the BSD benchmark and obtain results that are (in terms of accuracy) comparable to or better than those obtained with two popular spotfinder software packages ( Dozor and DIALS ), demonstrating that this is an appropriate framework to support future extensions and improvements.
Distributed energy resources and load management are an important emerging part of smart grids. Concurrent management of homeowner preferences and utility objectives requires an intelligent negotiation strategy supported by physical infrastructure that learns, optimizes, and controls the system. While there is a wide body of research related to modelling and simulation of distributed resources, relatively little is known about their practical feasibility. One of the first attempts to address this knowledge gap was through a 62-house connected neighborhood in Alabama. This study makes one more step in advancing this knowledge through a 46-townhome demonstration neighborhood located in Atlanta, GA. It reports hardware design, system architecture, and the results from the summer experimental work. The findings are discussed in the context of earlier experience, and analysis is expanded to areas which were not the focus of earlier research.
Aerosols and clouds are key components of the marine atmosphere, impacting the Earth’s radiative budget with a net cooling effect over the industrial era that counterbalances greenhouse gas warming, yet with an uncertain amplitude. Here we report recent advances in our understanding of how open ocean aerosol sources are modulated by ocean biogeochemistry and how they, in turn, shape cloud coverage and properties. We organize these findings in successive steps from ocean biogeochemical processes to particle formation by nucleation and sea spray emissions, further particle growth by condensation of gases, the potential to act as cloud condensation nuclei or ice nucleating particles, and finally, their effects on cloud formation, optical properties, and life cycle. We discuss how these processes may be impacted in a warming climate and the potential for ocean biogeochemistry—climate feedbacks through aerosols and clouds.
Breakthroughs from the U.S. Department of Energy Co-Optimization of Fuels & Engines (Co-Optima) initiative could make it possible to more rapidly cut emissions, reduce dependence on international petroleum, and contribute to ambitious national goals to slow global warming on the land, in the air, and across the water. Findings from the 6-year collaborative undertaking are now available in this report.
A myriad of phenomena in materials science and chemistry rely on quantum-level simulations of the electronic structure in matter. While moving to larger length and time scales has been a pressing issue for decades, such large-scale electronic structure calculations are still challenging despite modern software approaches and advances in high-performance computing. The silver lining in this regard is the use of machine learning to accelerate electronic structure calculations – this line of research has recently gained growing attention. The grand challenge therein is finding a suitable machine-learning model during a process called hyperparameter optimization. This, however, causes a massive computational overhead in addition to that of data generation. We accelerate the construction of machine-learning surrogate models by roughly two orders of magnitude by circumventing excessive training during the hyperparameter optimization phase. We demonstrate our workflow for Kohn-Sham density functional theory, the most popular computational method in materials science and chemistry.
This report summarizes the findings of the Geothermal Interagency Collaboration Task Force (Task Force) and associated stakeholder forums and Tribal listening sessions. The Task Force included federal agencies and state agencies in California and Nevada with a nexus to geothermal regulatory and permitting approvals. The Task Force met twice over the course of 2022 to discuss current geothermal regulatory and permitting challenges and strategies for improved coordination and permit processing. In addition, the project team held four forums/listening sessions in 2022 with geothermal industry representatives, environmental non-governmental organizations, and Tribes to gain additional insight and perspective on the geothermal regulatory process and managing potential cultural and natural resource conflicts that may arise during geothermal development.
This report summarizes a set of key findings that have been developed through a set of interconnected research activities performed by five institutions between January 2020 and December 2023. The project team, comprising Argonne National Laboratory, the National Renewable Energy Laboratory, Lawrence Berkeley National Laboratory, the Electric Power Research Institute, and Johns Hopkins University, collective engaged with the North American Independent System Operators and Regional Transmission Operators (ISO/RTOs) to identify the key challenges they are facing and opportunities for the project team to provide technical assistance in several prioritized challenge areas.
Random forests have become popular models used for data driven predictions. As a result, random forests are currently used or being considered for high-consequence mission applications in national security, such as the prediction of yield from optical signals and malware detection. While random forests may provide accurate predictions, the complexity of the algorithm causes a lack of interpretability. Random forests are an ensemble of regression or decision trees. Individual regression and decision trees are interpretable, but ensembles are inherently difficult to interpret due to the compilation of many models. We aim to increase the interpretability of random forests by finding patterns in the ensemble of trees that can be used to “thin” (or remove) trees. As a starting point, in this report, we develop a new distance metric for quantifying the similarity between trees based on their topologies (i.e., shapes). We base the metric on a novel distance metric for graphs that is a proper mathematical distance, is invariant to transformations, has registration between graphs, and computes topological evolutions between graphs. We use the tree distance metric to compute tree statistics such as a “mean tree” and to identify clusters of trees. We apply the developed methodology to a toy dataset and a mission relevant product inspection dataset to demonstrate how the metric can provide insight into random forests. Furthermore, we discuss the limitations of the approach and ideas for future research into how the metric could be used as a thinning tool to develop less complex models.
This report presents research findings from a four-year Smart Charge Management (SCM) pilot program conducted by Maryland’s largest electric utilities—Baltimore Gas and Electric (BGE), Potomac Electric Power Company (Pepco), and Delmarva Power & Light (DPL)—to evaluate strategies for optimizing electric vehicle (EV) charging loads and enhancing grid stability. Supported by the U.S. Department of Energy (DOE), Argonne National Laboratory collaborated with all project partners and examined the effectiveness of Time-of-Use (TOU) and Load Balancing (LB) strategies in managing peak demand, deferring costly infrastructure upgrades, and reducing grid constraints at the feeder level.
Under the Energy & Environmental Research Center’s (EERC’s) ~$\$$22 million Phase II Brine Extraction and Storage Test (BEST) Program, a multimillion-dollar brine treatment technology test bed facility was established in western North Dakota to provide a platform for evaluating developing technologies and approaches for brine treatment and volume reduction. The initial facility, and associated research effort, was funded by the U.S. Department of Energy (DOE), with in-kind contributions provided by several industry participants and the state of North Dakota. Since its opening, the Brine Technology Test Facility (BTTF) has supported performance evaluations of desalination technologies capable of treating high-salinity produced water (PW) and enabled data collection for multiple approaches of PW management and critical material recovery. As part of the decommissioning process for the original project, facility ownership and liability were transferred to Select Water, which is providing the EERC with a continuing site access option for state or federal research and/or commercial technology development. This report documents the findings from a design study conducted by the EERC and the engineering firm that was originally contracted to design and construct the facility (Advanced Engineering and Environmental Services, LLC [AE2S]) that evaluated the current status of BTTF and its systems and developed a retrofit design to increase the facility’s capabilities through automation and modularization of its PW treatment infrastructure. The proposed retrofit will provide DOE and industry with an expanded range of conditioned PW that can be produced at the facility for evaluating fit-for-purpose water treatment technologies, online instrumentation for brine chemistry determination, and systems that recover critical materials like lithium and magnesium. The current facility consists of an 9600-square-foot facility that includes a 40-foot by 65-foot Class 1, Division 2-rated demonstration area and associated control rooms and lab-ready space capable of sourcing oil and gas PW and wastewater from industrial sources or tailoring brine compositions up to 300,000 mg/L total dissolved solids (TDS) and supplying them at rates up to 25 gpm for extended-duration technology demonstrations. The colocation of the facility with Select Water’s water management facilities allows for access and unloading of more than 10,000 bbl/day of trucked water delivered to site and associated access to on-site Class I and Class II brine disposal wells and nearby hazardous waste landfills operated and/or contracted by Select Water to dispose of concentrate and/or effluents associated with the testing.