Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Jupyter”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Tracking atomic structure evolution during directed electron beam induced Si-atom motion in graphene via deep machine learning

Using electron beam manipulation, we enable deterministic motion of individual Si atoms in graphene along predefined trajectories. Structural evolution during the dopant motion was explored, providing information on changes of the Si atom neighborhood during atomic motion and providing statistical information of possible defect configurations. The combination of a Gaussian mixture model and principal component analysis applied to the deep learning-processed experimental data allowed disentangling of the atomic distortions for two different graphene sublattices. This approach demonstrates the potential of e-beam manipulation to create defect libraries of multiple realizations of the same defect and explore the potential of symmetry breaking physics. The rapid image analytics enabled via a deep learning network further empowers instrumentation for e-beam controlled atom-by-atom fabrication. Here, the analysis described in the paper can be reproduced via an interactive Jupyter notebook at https://git.io/JJ3Bx.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

End-to-end AI framework for interpretable prediction of molecular and crystal properties

We introduce an end-to-end computational framework that allows for hyperparameter optimization using the DeepHyper library, accelerated model training, and interpretable AI inference. The framework is based on state-of-the-art AI models including CGCNN, PhysNet, SchNet, MPNN, MPNN-transformer, and TorchMD-NET. We employ these AI models along with the benchmark QM9, hMOF, and MD17 datasets to showcase how the models can predict user-specified material properties within modern computing environments. We demonstrate transferable applications in the modeling of small molecules, inorganic crystals and nanoporous metal organic frameworks with a unified, standalone framework. We have deployed and tested this framework in the ThetaGPU supercomputer at the Argonne Leadership Computing Facility, and in the Delta supercomputer at the National Center for Supercomputing Applications to provide researchers with modern tools to conduct accelerated AI-driven discovery in leadership-class computing environments. We release these digital assets as open source scientific software in GitLab, and ready-to-use Jupyter notebooks in Google Colab.

36 MATERIALS SCIENCE↗

Deep kernel methods learn better: from cards to process optimization

Abstract The ability of deep learning methods to perform classification and regression tasks relies heavily on their capacity to uncover manifolds in high-dimensional data spaces and project them into low-dimensional representation spaces. In this study, we investigate the structure and character of the manifolds generated by classical variational autoencoder (VAE) approaches and deep kernel learning (DKL). In the former case, the structure of the latent space is determined by the properties of the input data alone, while in the latter, the latent manifold forms as a result of an active learning process that balances the data distribution and target functionalities. We show that DKL with active learning can produce a more compact and smooth latent space which is more conducive to optimization compared to previously reported methods, such as the VAE. We demonstrate this behavior using a simple cards dataset and extend it to the optimization of domain-generated trajectories in physical systems. Our findings suggest that latent manifolds constructed through active learning have a more beneficial structure for optimization problems, especially in feature-rich target-poor scenarios that are common in domain sciences, such as materials synthesis, energy storage, and molecular discovery. The Jupyter Notebooks that encapsulate the complete analysis accompany the article.

97 MATHEMATICS AND COMPUTING↗

A mixed, unified forward/inverse framework for earthquake problems: fault implementation and coseismic slip estimate

SUMMARY We introduce a new finite-element (FE) based computational framework to solve forward and inverse elastic deformation problems for earthquake faulting via the adjoint method. Based on two advanced computational libraries, FEniCS and hIPPYlib for the forward and inverse problems, respectively, this framework is flexible, transparent and easily extensible. We represent a fault discontinuity through a mixed FE elasticity formulation, which approximates the stress with higher order accuracy and exposes the prescribed slip explicitly in the variational form without using conventional split node and decomposition discrete approaches. This also allows the first order optimality condition, that is the vanishing of the gradient, to be expressed in continuous form, which leads to consistent discretizations of all field variables, including the slip. We show comparisons with the standard, pure displacement formulation and a model containing an in-plane mode II crack, whose slip is prescribed via the split node technique. We demonstrate the potential of this new computational framework by performing a linear coseismic slip inversion through adjoint-based optimization methods, without requiring computation of elastic Green’s functions. Specifically, we consider a penalized least squares formulation, which in a Bayesian setting—under the assumption of Gaussian noise and prior—reflects the negative log of the posterior distribution. The comparison of the inversion results with a standard, linear inverse theory approach based on Okada’s solutions shows analogous results. Preliminary uncertainties are estimated via eigenvalue analysis of the Hessian of the penalized least squares objective function. Our implementation is fully open-source and Jupyter notebooks to reproduce our results are provided. The extension to a fully Bayesian framework for detailed uncertainty quantification and non-linear inversions, including for heterogeneous media earthquake problems, will be analysed in a forthcoming paper.

58 GEOSCIENCES↗

CLMM : a LSST-DESC cluster weak lensing mass modeling library for cosmology

ABSTRACT We present the v1.0 release of CLMM, an open source python library for the estimation of the weak lensing masses of clusters of galaxies. CLMM is designed as a stand-alone toolkit of building blocks to enable end-to-end analysis pipeline validation for upcoming cluster cosmology analyses such as the ones that will be performed by the Vera C. Rubin Legacy Survey of Space and Time-Dark Energy Science Collaboration (LSST-DESC). Its purpose is to serve as a flexible, easy-to-install, and easy-to-use interface for both weak lensing simulators and observers and can be applied to real and mock data to study the systematics affecting weak lensing mass reconstruction. At the core of CLMM are routines to model the weak lensing shear signal given the underlying mass distribution of galaxy clusters and a set of data operations to prepare the corresponding data vectors. The theoretical predictions rely on existing software, used as backends in the code, that have been thoroughly tested and cross-checked. Combined theoretical predictions and data can be used to constrain the mass distribution of galaxy clusters as demonstrated in a suite of example Jupyter Notebooks shipped with the software and also available in the extensive online documentation.

79 ASTRONOMY AND ASTROPHYSICS↗

ClusterCAD 2.0: an updated computational platform for chimeric type I polyketide synthase and nonribosomal peptide synthetase design

Abstract Megasynthase enzymes such as type I modular polyketide synthases (PKSs) and nonribosomal peptide synthetases (NRPSs) play a central role in microbial chemical warfare because they can evolve rapidly by shuffling parts (catalytic domains) to produce novel chemicals. If we can understand the design rules to reshuffle these parts, PKSs and NRPSs will provide a systematic and modular way to synthesize millions of molecules including pharmaceuticals, biomaterials, and biofuels. However, PKS and NRPS engineering remains difficult due to a limited understanding of the determinants of PKS and NRPS fold and function. We developed ClusterCAD to streamline and simplify the process of designing and testing engineered PKS variants. Here, we present the highly improved ClusterCAD 2.0 release, available at https://clustercad.jbei.org. ClusterCAD 2.0 boasts support for PKS-NRPS hybrid and NRPS clusters in addition to PKS clusters; a vastly enlarged database of curated PKS, PKS-NRPS hybrid, and NRPS clusters; a diverse set of chemical ‘starters’ and loading modules; the new Domain Architecture Cluster Search Tool; and an offline Jupyter Notebook workspace, among other improvements. Together these features massively expand the chemical space that can be accessed by enzymes engineered with ClusterCAD.

59 BASIC BIOLOGICAL SCIENCES↗

Quantifying uncertainties and correlations in the nuclear-matter equation of state

We perform statistically rigorous uncertainty quantification (UQ) for chiral effective field theory (χ EFT) applied to infinite nuclear matter up to twice nuclear saturation density. The equation of state (EOS) is based on high-order many-body perturbation theory calculations with nucleon-nucleon and three-nucleon interactions up to fourth order in the χ EFT expansion. From these calculations our newly developed Bayesian machine-learning approach extracts the size and smoothness properties of the correlated EFT truncation error. Furthermore, we then propose a novel extension that uses multitask machine learning to reveal correlations between the EOS at different proton fractions. The inferred in-medium χ EFT breakdown scale in pure neutron matter and symmetric nuclear matter is consistent with that from free-space nucleon-nucleon scattering. These significant advances allow us to provide posterior distributions for the nuclear saturation point and propagate theoretical uncertainties to derived quantities: the pressure and incompressibility of symmetric nuclear matter, the nuclear symmetry energy, and its derivative. Our results, which are validated by statistical diagnostics, demonstrate that an understanding of truncation-error correlations between different densities and different observables is crucial for reliable UQ. The methods developed here are publicly available as annotated Jupyter notebooks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Wave-function-based emulation for nucleon-nucleon scattering in momentum space

Emulators for low-energy nuclear physics can provide fast and accurate predictions of bound-state and scattering observables for applications that require repeated calculations with different parameters, such as Bayesian uncertainty quantification. In this paper, we extend a scattering emulator based on the Kohn variational principle (KVP) to momentum space (including coupled channels) with arbitrary boundary conditions, which enable the mitigation of spurious singularities known as Kohn anomalies. We test it on a modern chiral nucleon-nucleon (N N) interaction, including emulation of the coupled channels. We provide comparisons between a Lippmann-Schwinger equation emulator and our KVP momentum-space emulator for a representative set of neutron-proton (n p) scattering observables, and also introduce a quasi-spline-based approach for the KVP-based emulator. Furthermore, our findings show that while there are some trade-offs between accuracy and speed, all three emulators perform well. Self-contained Jupyter notebooks that generate the results and figures in this paper are publicly available.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Assessing correlated truncation errors in modern nucleon-nucleon potentials

We test the BUQEYE model of correlated effective field theory (EFT) truncation errors on Reinert, Krebs, and Epelbaum's semilocal momentum-space implementation of the chiral EFT (𝜒⁢EFT ) expansion of the nucleon-nucleon (NN) potential. This Bayesian model hypothesizes that dimensionless coefficient functions extracted from the order-by-order corrections to NN observables can be treated as draws from a Gaussian process (GP). We combine a variety of graphical and statistical diagnostics to assess when predicted observables have a 𝜒⁢EFT convergence pattern consistent with the hypothesized GP statistical model. Our conclusions are that, first, the BUQEYE model is generally applicable to the potential investigated here, which enables statistically principled estimates of the impact of higher EFT orders on observables. Second, parameters defining the extracted coefficients such as the expansion parameter 𝑄 must be well chosen for the coefficients to exhibit a regular convergence pattern—a property we exploit to obtain posterior distributions for such quantities. Third, the assumption of GP stationarity across lab energy and scattering angle is not generally met; this necessitates adjustments in future work. We provide a workflow and interpretive guide for our analysis framework, and show what can be inferred about probability distributions for 𝑄, the EFT breakdown scale Λ 𝑏 , the scale associated with soft physics in the 𝜒⁢EFT potential 𝑚 eff , and the GP hyperparameters. All our results can be reproduced using a publicly available Jupyter notebook, which can be straightforwardly modified to analyze other 𝜒⁢EFT NN potentials.

Bayesian methods↗

Addition of tabulated equation of state and neutrino leakage support to illinoisgrmhd

Here we have added support for realistic, microphysical, finite-temperature equations of state (EOS) and neutrino physics via a leakage scheme to illinoisgrmhd, an open-source GRMHD code for dynamical spacetimes in the einstein toolkit. These new features are provided by two new, nrpy+-based codes: nrpyeos, which performs highly efficient EOS table lookups and interpolations, and nrpyleakage, which implements a new, adaptive mesh refinement (AMR)-capable neutrino leakage scheme in the einstein toolkit. We have performed a series of strenuous validation tests that demonstrate the robustness of these new codes, particularly on the Cartesian AMR grids provided by carpet. Furthermore, we show results from fully dynamical GRMHD simulations of single unmagnetized neutron stars, and magnetized binary neutron star mergers. This new version of illinoisgrmhd, as well as nrpyeos and nrpyleakage, is pedagogically documented in jupyter notebooks and fully open source. The codes will be proposed for inclusion in an upcoming version of the einstein toolkit.

79 ASTRONOMY AND ASTROPHYSICS↗

Solution scattering at the Life Science X-ray Scattering (LiX) beamline

This work reports the instrumentation and software implementation at the Life Science X-ray Scattering (LiX) beamline at NSLS-II in support of biomolecular solution scattering. For automated static measurements, samples are stored in PCR tubes and grouped in 18-position sample holders. Unattended operations are enabled using a six-axis robot that exchanges sample holders between a storage box and a sample handler, transporting samples from the PCR tubes to the X-ray beam for scattering measurements. The storage box has a capacity of 20 sample holders. At full capacity, the measurements on all samples last for ∼9 h. For in-line size-exclusion chromatography, the beamline-control software coordinates with a commercial high-performance liquid chromatography (HPLC) system to measure multiple samples in batch mode. The beamline can switch between static and HPLC measurements instantaneously. In all measurements, the scattering data span a wide q -range of typically 0.006–3.2 Å −1 . Functionalities in the Python package py4xs have been developed to support automated data processing, including azimuthal averaging, merging data from multiple detectors, buffer scattering subtraction, data storage in HDF5 format and exporting the final data in a three-column text format that is acceptable by most data analysis tools. These functionalities have been integrated into graphical user interfaces that run in Jupyter notebooks, with hooks for external data analysis software.

60 APPLIED LIFE SCIENCES↗

tomoCAM : fast model-based iterative reconstruction via GPU acceleration and non-uniform fast Fourier transforms

X-ray-based computed tomography is a well established technique for determining the three-dimensional structure of an object from its two-dimensional projections. In the past few decades, there have been significant advancements in the brightness and detector technology of tomography instruments at synchrotron sources. These advancements have led to the emergence of new observations and discoveries, with improved capabilities such as faster frame rates, larger fields of view, higher resolution and higher dimensionality. These advancements have enabled the material science community to expand the scope of tomographic measurements towards increasingly in situ and in operando measurements. In these new experiments, samples can be rapidly evolving, have complex geometries and restrictions on the field of view, limiting the number of projections that can be collected. In such cases, standard filtered back-projection often results in poor quality reconstructions. Iterative reconstruction algorithms, such as model-based iterative reconstructions (MBIR), have demonstrated considerable success in producing high-quality reconstructions under such restrictions, but typically require high-performance computing resources with hundreds of compute nodes to solve the problem in a reasonable time. Here, tomoCAM , is introduced, a new GPU-accelerated implementation of model-based iterative reconstruction that leverages non-uniform fast Fourier transforms to efficiently compute Radon and back-projection operators and asynchronous memory transfers to maximize the throughput to the GPU memory. The resulting code is significantly faster than traditional MBIR codes and delivers the reconstructive improvement offered by MBIR with affordable computing time and resources. tomoCAM has a Python front-end, allowing access from Jupyter -based frameworks, providing straightforward integration into existing workflows at synchrotron facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Normality of I-V Measurements Using ML

There is an increased interest in instrument-computing ecosystems (ICEs) that support science workflows empowered by AI-automated experiments and computations in diverse areas. In particular, electrochemistry ICEs are promising for accelerating the design and discovery of electrochemical systems for energy storage and conversion, by automating significant parts of workflows that combine synthesis and characterization experiments with computations. They require the integration of flow controllers, solvent containers, pumps, fraction collectors, and potentiostats, all connected to an electrochemical cell, as illustrated in Fig. 1. These are specialized instruments with custom software that is not originally designed for network integration. We developed network and software solutions for electrochemical workflows that adapt system and instrument settings in real-time for multiple rounds of experiments. In particular, we developed Python wrappers for Application Programming Interfaces (APIs) of instrument commands and Pyro client-server modules that enable them to be executed from remote computers. The entire workflow is orchestrated by a Jupyter notebook running on a remote computer.

Al Najjar, Anees↗

Descriptor: High Temporal Resolution Meteorological Data at Oak Ridge Reservation (ORR-HiResMet)

Access to continuous, quality assessed meteorological data is critical for understanding the climatology and atmospheric dynamics of a region. Research facilities like Oak Ridge National Laboratory (ORNL) rely on such data to assess site-specific climatology, model potential emissions, establish safety baselines, and prepare for emergency scenarios. To meet these needs, on-site towers at ORNL collect meteorological data at 15-minute and hourly intervals. However, data measurements from meteorological towers are affected by sensor sensitivity, degradation, lightning strikes, power fluctuations, glitching, and sensor failures, all of which can affect data quality. To address these challenges, we conducted a comprehensive quality assessment and processing of five years of meteorological data collected from ORNL at 15-minute intervals, including measurements of temperature, pressure, humidity, wind, and solar radiation. The time series of each variable was pre-processed and gap-filled using established meteorological data collection and cleaning techniques, i.e., the time series were subjected to structural standardization, data integrity testing, automated and manual outlier detection, and gap-filling. The data product and highly generalizable processing workflow developed in Python Jupyter notebooks are publicly accessible online. As a key contribution of this study, the evaluated 5-year data will be used to train atmospheric dispersion models that simulate dispersion dynamics across the complex ridge-and-valley topography of the Oak Ridge Reservation in East Tennessee.

Steckler, Morgan R. [Oak Ridge National Laboratory↗

Virtual Framework for Development and Testing of Federation Software Stack

Softwarization of networked infrastructures combined with containerization of codes promises unprecedented computing capabilities distributed across the federations of computing systems and physical instruments. The development and testing of a software stack that implements these capabilities over an expensive physical production infrastructure is not cost-effective, and in the early stages, may potentially cause service disruptions. To address these aspects, we develop the Virtual Federated Science Instrument Environment (VFSIE), a digital twin of the physical infrastructure that emulates a multi-site federation. Each federated site is emulated using containers and virtual hosts that are connected over local-area networks, and the sites, in turn, are connected over an emulated wide-area network. We describe the framework design and implementation details. We also illustrate its application by emulating a federation of four laboratories that use Jupyter Notebook for computations and the EPICS software system for instrument control.

Al Najjar, Anees↗

Modification and analysis of context-specific genome-scale metabolic models: methane-utilizing microbial chassis as a case study

ABSTRACT Context-specific genome-scale model (CS-GSM) reconstruction is becoming an efficient strategy for integrating and cross-comparing experimental multi-scale data to explore the relationship between cellular genotypes, facilitating fundamental or applied research discoveries. However, the application of CS modeling for non-conventional microbes is still challenging. Here, we present a graphical user interface that integrates COBRApy, EscherPy, and RIPTiDe, Python-based tools within the BioUML platform, and streamlines the reconstruction and interrogation of the CS genome-scale metabolic frameworks via Jupyter Notebook. The approach was tested using -omics data collected for Methylotuvimicrobium alcaliphilum 20Z R , a prominent microbial chassis for methane capturing and valorization. We optimized the previously reconstructed whole genome-scale metabolic network by adjusting the flux distribution using gene expression data. The outputs of the automatically reconstructed CS metabolic network were comparable to manually optimized i IA409 models for Ca-growth conditions. However, the CS model questions the reversibility of the phosphoketolase pathway and suggests higher flux via primary oxidation pathways. The model also highlighted unresolved carbon partitioning between assimilatory and catabolic pathways at the formaldehyde-formate node. Only a very few genes and only one enzyme with a predicted function in C1 metabolism, a homolog of the formaldehyde oxidation enzyme ( fae1-2 ), showed a significant change in expression in La-growth conditions. The CS-GSM predictions agreed with the experimental measurements under the assumption that the Fae1-2 is a part of the tetrahydrofolate-linked pathway. The cellular roles of the tungsten (W)-dependent formate dehydrogenase ( fdhAB ) and fae homologs ( fae1-2 and fae3 ) were investigated via mutagenesis. The phenotype of the f dhAB mutant followed the model prediction. Furthermore, a more significant reduction of the biomass yield was observed during growth in La-supplemented media, confirming a higher flux through formate. M. alcaliphilum 20Z R mutants lacking fae1-2 did not display any significant defects in methane or methanol-dependent growth. However, contrary to fae1, the fae1-2 homolog failed to restore the formaldehyde-activating enzyme function in complementation tests. Overall, the presented data suggest that the developed computational workflow supports the reconstruction and validation of CS-GSM networks of non-model microbes. IMPORTANCE The interrogation of various types of data is a routine strategy to explore the relationship between genotype and phenotype. An efficient approach for integrating and cross-comparing experimental multi-scale data in the context of whole-genome-based metabolic network reconstruction becomes a powerful tool that facilitates fundamental and applied research discoveries. The present study describes the reconstruction of a context-specific (CS) model for the methane-utilizing bacterium, Methylotuvimicrobium alcaliphilum 20Z R . M. alcaliphilum 20Z R is becoming an attractive microbial platform for the production of biofuels, chemicals, pharmaceuticals, and bio-sorbents for capturing atmospheric methane. We demonstrate that this pipeline can help reconstruct metabolic models that are similar to manually curated networks. Furthermore, the model is able to highlight previously overlooked pathways, thus advancing fundamental knowledge of non-model microbial systems or promoting their development toward biotechnological or environmental implementations.

Kulyashov, M. A.↗

Quantifying the Impact of Advanced Web Platforms on High Performance Computing Usage

The deployment of Science Gateways for High Performance Computing (HPC) systems can alter long-accepted usage patterns on supercomputing systems in positive ways as an ever-increasing number of users migrate their workflows to HPC systems. Idaho National Laboratory (INL) has deployed two separate advanced web platforms, Open OnDemand and NICE DCV, for integration with HPC resources to improve web accessibility for HPC users. Researchers conducted a multi-year study on how HPC usage pat- terns changed in the presence of these platforms. This work reports the results of that study and quantifies the observed impacts, including adoption by visualization and Jupyter Notebook/Lab users, decreased job submission friction, rapid uptake of HPC by Windows users, and increased overall system utilization. The most significant impacts were observed from the deployment of Open OnDemand, and this work also identifies some best practices for Open OnDemand deployment for HPC datacenters.

97 MATHEMATICS AND COMPUTING↗

Julia as a unifying end-to-end workflow language on the Frontier exascale system

We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point stencil kernel on Frontier, the US Department of Energy’s first exascale supercomputer. We evaluate the performance, scaling, and trade-offs of (i) the computational kernel on AMD’s MI250x GPUs, (ii) weak scaling up to 4,096 MPI processes/GPUs or 512 nodes, (iii) parallel I/O writes using the ADIOS2 library bindings, and (iv) Jupyter Notebooks for interactive analysis. Results suggest that although Julia generates a reasonable LLVM-IR, a nearly 50% performance difference exists vs. native AMD HIP stencil codes when running on the GPUs. As expected, we observed near-zero overhead when using MPI and parallel I/O bindings for system-wide installed implementations. Consequently, Julia emerges as a compelling high-performance and high-productivity workflow composition language, as measured on the fastest supercomputer in the world.

Godoy, William↗