Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Python codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Mshpy23: A User-Friendly, Parameterized Model of Magnetosheath Conditions

Lunar Environment heliospheric X-ray Imager (LEXI) and Solar wind – Magnetosphere - Ionosphere Link Explorer (SMILE) will observe magnetosheath and its boundary motion in soft X-rays for understanding magnetopause reconnection modes under varioussolar wind conditions after their respective launches in 2024 and 2025. Magnetosheath conditions, namely, plasma density, velocity, and temperature, are key parameters for predicting and analyzing soft X-ray images from the LEXI and SMILE missions. We developed a user-friendly model of magnetosheath that parameterizes number density, velocity, temperature, and magnetic field by utilizing the global Magnetohydrodynamics (MHD) model as well as the pre-existing gas-dynamic and analytic models. Using this parameterized magnetosheath model, scientists can easily reconstruct expected soft X-ray images and utilize them for analysis of observed images of LEXI and SMILE without simulating the complicated global magnetosphere models. First, we created an MHD based magnetosheath model by running a total of 14 OpenGGCM global MHD simulations under 7 solar wind densities (1, 5, 10, 15, 20, 25, and 30cm−3) and 2 interplanetary magnetic field BZ components (± 4nT), and then parameterizing the results in new magnetosheath conditions. We compared the magnetosheath model result with THEMIS statistical data and it showed good agreement with a weighted Pearson correlation coefficient greater than 0.77, especially for plasma density and plasma velocity. Second, we compiled a suite of magnetosheath models incorporating previous magnetosheath models (gas-dynamic, analytic), and did two case studies to test the performance. The MHD based model was comparable to or better than the previous models while providing self consistency among the magnetosheath parameters. Third, we constructed a tool to calculate a soft X-ray image from any given vantage point, which can support the planning and data analysis of the aforementioned LEXI and SMILE missions. A release of the code has been uploaded to a Github repository.

Magnetosheath↗

PyCDFT: A Python package for constrained density functional theory

In this paper, we present PyCDFT, a Python package to compute diabatic states using constrained density functional theory (CDFT). PyCDFT provides an object-oriented, customizable implementation of CDFT, and allows for both single-point self-consistent-field calculations and geometry optimizations. PyCDFT is designed to interface with existing density functional theory (DFT) codes to perform CDFT calculations where constraint potentials are added to the Kohn–Sham Hamiltonian. Here, we demonstrate the use of PyCDFT by performing calculations with a massively parallel first-principles molecular dynamics code, Qbox, and we benchmark its accuracy by computing the electronic coupling between diabatic states for a set of organic molecules. We show that PyCDFT yields results in agreement with existing implementations and is a robust and flexible package for performing CDFT calculations. The program is available at https://dx.doi.org/10.5281/zenodo.3821097.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Kimberlina 1.2 CCUS Geophysical Models and Synthetic Data Sets

This synthetic multi-scale and multi-physics data set was produced in collaboration with teams at the Lawrence Berkeley National Laboratory, National Energy Technology Laboratory, Los Alamos National Laboratory, and Colorado School of Mines through the Science-informed Machine Learning for Accelerating Real-Time Decisions in Subsurface Applications (SMART) Initiative. Data are associated with the following publication: Alumbaugh, D., Gasperikova, E., Crandall, D., Commer, M., Feng, S., Harbert, W., Li, Y., Lin, Y., and Samarasinghe, S., “The Kimberlina Synthetic Geophysical Model and Data Set for CO2 Monitoring Investigations”, The Geoscience Data Journal, 2023, DOI: 10.1002/gdj3.191. The dataset uses the Kimberlina 1.2 CO2 reservoir flow model simulations based on a hypothetical CO2 storage site in California (Birkholzer et al., 2011; Wainwright et al., 2013). Geophysical properties models (P- and S-wave seismic velocities, saturated density, and electrical resistivity) were produced with an approach similar to that of Yang et al. (2019) and Gasperikova et al. (2022) for 100 Kimberlina 1.2 reservoir models. Links to individual resources are provided below: [CO2 Saturation Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-co2-saturation-models); Resistivity Models – [part 1](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-1), [part 2](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-2), and [part 3](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-resistivity-models-part-3); [Vp Velocity Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-vp-velocity-models); [Vs Velocity Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-vs-velocity-models); [Density Models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-density-models). The 3D distributions of geophysical properties for the 33 time stamps of the SIM001 model were used to generate synthetic seismic, gravity, and electromagnetic (EM) responses for 33 times between zero and 200 years. Synthetic surface seismic data were generated using 2D and 3D finite-difference codes that simulate the acoustic wave equation (Moczo et al., 2007). 2D data were simulated for six point-pressure sources along a 2D line with 10 m receiver spacing and a time spacing of 0.0005 s. 3D simulations were completed for 25 surface pressure sources using a source separation of 1 km in both the x and y directions and a time spacing of 0.001 s. Links to individual resources are provided below: [2D velocity models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-2d-velocity-models) and [2D surface seismic data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-2d-surface-seismic-data). [3D velocity models](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-velocity-models), and 3D seismic data [year0](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year0), [year1](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year1), [year2](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year2), [year5](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year5), [year10](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year10), [year15](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year15), [year20](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year20), [year25](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year25), [year30](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year30), [year35](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year35), [year40](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year40), [year45](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year45), [year49](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year49), [year50](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year50), [year51](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year51), [year52](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year52), [year55](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year55), [year60](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year60), [year65](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year65), [year70](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year70), [year75](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year75), [year80](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year80), [year85](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year85), [year90](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year90), [year95](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year95), [year100](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year100), [year110](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year110), [year120](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year120), [year130](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year130), [year140](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year140), [year150](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year150), [year175](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year175), [year200](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-seismic-data-year200). The Python scripts to read these models and data are provided [here](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-python-scripts). EM simulations used a borehole-to-surface survey configuration, with the source located near the reservoir level and receivers on the surface using the code developed by Commer and Newman (2008). Pseudo-2D data for the source at [2500 m](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-pseudo-2d-csem-data-tz2500m) and [3025 m](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-pseudo-2d-csem-data-tz3025m), used a 2D inline receiver configuration to simulate a response over 3D resistivity models. The [3D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-3d-csem-data) contain electric fields generated by borehole sources at monitoring well locations and measured over a surface receiver grid. Vector gravity data, both on the surface and in boreholes, were simulated using a modeling code developed by Rim and Li (2015). The simulation scenarios were parallel to those used for the EM: [pseudo-2D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-gravity-data) were calculated along the same lines and within the same boreholes, and [3D data](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-gravity-data) were simulated over 3D models on the surface and in three monitoring wells. A series of [synthetic well logs](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-well-logs) of CO2 saturation, acoustic velocity, density, and induction resistivity in the injection well and three monitoring wells are also provided at 0, 1, 2, 5, 10, 15, and 20 years after the initiation of injection. These were constructed by combining the low-frequency trend of the geophysical models with the high-frequency variations of actual well logs collected in the Kimberlina 1 well that was drilled at the proposed site. Measurements of permeability and pore connectivity were made on cores of Vedder Sandstone, which forms the primary reservoir unit: [CT micro scans](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-ct-micro-scans-of-vedder-formation) and [Industrial CT Images](https://edx.netl.doe.gov/dataset/kimberlina-1-2-ccus-geophysical-models-and-synthetic-data-sets-industrial-ct-images-vedder-formation). These measurements provide the range of scales in the otherwise synthetic data set to be as close to a real-world situation as possible. References: Birkholzer, J.T., Zhou, Q., Cortis, A. and Finsterle, S., 2011. A sensitivity study on regional pressure buildup from large-scale CO2 storage projects. Energy Procedia, 4, 4371-4378. Commer, M., and Newman, G.A., 2008. New advances in three-dimensional controlled-source electromagnetic inversion, Geophysical Journal International, 172, 513-535. Gasperikova, E., Appriou, D., Bonneville, A., Feng, Z., Huang, L., Gao, K., Yang, X., Daley, T., 2022, Sensitivity of geophysical techniques for monitoring secondary CO2 storage plumes, Int. J. Greenh. Gas Control, Volume 114, 103585, ISSN 1750-5836, https://doi.org/10.1016/j.ijggc.2022.103585. Moczo, P., J.O. Robertsson and L. Eisner, 2007, The finite-difference time-domain method for modeling of seismic wave propagation: Advances in geophysics, 48, 421-516. Rim, H., and Y. Li, 2015, Advantages of borehole vector gravity in density imaging, Geophysics, 80, G1-G13. Wainwright, H. M.; Finsterle, S.; Zhou, Q.; Birkholzer, J. T., 2013. Modeling the Performance of Large-Scale CO2 Storage Systems: A Comparison of Different Sensitivity Analysis Methods. International Journal of Greenhouse Gas Control, 17, 189205. https://doi.org/10.1016/j.ijggc.2013.05.007, DOI: 10.18141/1603331. Yang, X., Buscheck, T.A., Mansoor, K., Wang, Z., Gao, K., Huang, L., Appriou, D., and Carroll, S.A., 2019. Assessment of geophysical monitoring methods for detection of brine and CO2 leakage in drinking water aquifers, International Journal of Greenhouse Gas Control, 90, 102803, https://doi.org/10.1016/j.ijggc.2019.102803.

CCUS↗

Language Independent Static Analysis (LISA)

Software is becoming increasingly important in nearly every aspect of global society and therefore in nearly every aspect of national security as well. While there have been major advancements in recent years in formally proving properties of program source code during development, such approaches are still in the minority among development teams, and the vast majority of code in this software explosion is produced without such properties. In these cases, the source code must be analyzed in order to establish whether the properties of interest hold. Because of the volume of software being produced, automated approaches to software analysis are necessary to meet the need. However, this software boom is not occurring in just one language. There are a wide range of languages of interest in national security spaces, including well-known languages such as C, C++, Python, Java, Javascript, and many more. But recent years have produced a wide range of new languages, including Nim, (2008), Go (2009), Rust (2010), Dart (2011), Kotlin (2011), Elixir (2011), Red (2011), Julia (2012), Typescript (2012), Swift (2014), Hack (2014), Crystal (2014), Ballerina (2017) and more. Historically, automated software analyses are implemented as tools that intermingle both the analysis question at hand with target language dependencies throughout their code, making re-use of components for different analysis questions or different target languages impractical. This project seeks to explore how mission-relevant, static software analyses can be designed and constructed in a language-independent fashion, dramatically increasing the reusability of software analysis investments.

97 MATHEMATICS AND COMPUTING↗

Evaluating Awkward Arrays, uproot, and coffea as a query platform for High Energy Physics Data

Query languages for High Energy Physics (HEP) are an ever present topic within the field. A query language that can efficiently represent the nested data structures that encode the statistical and physical meaning of HEP data will help analysts by ensuring their code is more clear and pertinent. As the result of a multi-year effort to develop an in-memory columnar representation of high energy physics data, the NumPy, Awkward Array, and uproot Python packages present a mature and efficient interface to HEP data. Atop that base, the coffea package adds functionality to launch queries at scale, manage and apply experiment-specific transformations to data, and present a rich object-oriented columnar data representation to the analyst. Recently, a set of Analysis Description Language (ADL) benchmarks has been established to compare HEP queries in multiple languages and frameworks. In this paper we present these benchmark queries implemented within the coffea framework and discuss their readability and performance characteristics. We find that the columnar queries perform as well or better than the implementations given in previous studies.

Gray, L.↗

A software package for plasma facing component analysis and design: the Heat flux Engineering Analysis Toolkit (HEAT)

The engineering limits of plasma facing components (PFCs) constrain the allowable operational space of tokamaks. Poorly managed heat fluxes that push the PFCs beyond their limits not only degrade core plasma performance via elevated impurities, but can also result in PFC failure due to thermal stresses or melting. Simple axisymmetric assumptions fail to capture the complex interaction between 3D PFC geometry and 2D or 3D plasmas. This results in fusion systems that must either operate with increased risk or reduce PFC loads, potentially through lower core plasma performance, to maintain a nominal safety factor. High precision 3D heat flux predictions are necessary to accurately ascertain the state of a PFC given the evolution of the magnetic equilibrium. A new code, the Heat flux Engineering Analysis Toolkit (HEAT), has been developed to provide high precision 3D predictions and analysis for PFCs. HEAT couples many otherwise disparate computational tools together into a single open source python package. Magnetic equilibrium, engineering CAD, finite volume solvers, scrape off layer plasma physics, visualization, high performace computing, and more, are connected in a single web-based user interface. Linux users may use HEAT without any software prerequisites via an appImage. This manuscript introduces HEAT, discusses the software architecture, presents first HEAT results, and outlines physics modules in development.

divertor physics↗

PYOED: AN ETENSIBLE SUITE FOR DATA ASSIMILATION AND MODEL-CONSTRAINED OPTIMAL DESIGN OF EXPERIMENTS

SF-23-005 PyOED is a highly extensible scientific package that enables developing and testing model-constrained optimal experimental design (OED) for inverse problems. Specifically, PyOED aims to be a comprehensive Python toolkit for model-constrained OED. The package targets scientists and researchers interested in understanding the details of OED formulations and approaches. It is also meant to enable researchers to experiment with standard and innovative OED technologies with a wide range of test problems (e.g., simulation models). OED, inverse problems (e.g., Bayesian inversion), and data assimilation (DA) are closely related research fields, and their formulations overlap significantly. Thus, PyOED is continuously being expanded with a plethora of Bayesian inversion, DA, and OED methods as well as new scientific simulation models, observation error models, and observation operators. These pieces are added such that they can be permuted to enable testing OED methods in various settings of varying complexities. The PyOED core is completely written in Python and utilizes the inherent object-oriented capabilities; however, PyOED is meant to be extensible rather than scalable. Specifically, PyOED is developed to ``enable rapid development and benchmarking of OED methods with minimal coding effort and to maximize code reutilization.'' PyOED will be continuously expanded with a plethora of Bayesian inversion, DA, and OED methods as well as new scientific simulation models, observation error models, and observation operators.

ATTIA, AHMEDMOHAMED↗

HPB_strengthmodel

Python-implementation of the Hunter-Preston strength model, but with generalized drag coefficient B. For details, see D. N. Blaschke, A. Hunter, and D. L. Preston, Int. J. Plast. 131 (2020) 102750. This code was used to generate most of the figures in this paper.

Blaschke, Daniel N.↗

National Climate Database (NCDB)

The National Climate Database (NCDB) is a high resolution, bias-corrected climate dataset consisting of the three most widely used variables of solar radiation- global horizontal (GHI), direct normal (DNI), and diffuse horizontal irradiance (DHI)- as well as other meteorological data. The goal of the NCDB is to provide unbiased high temporal and spatial resolution climate data needed for renewable energy modeling. The NCDB is modeled using a statistical downscaling approach with Regional Climate Model (RCM)-based climate projections obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX; linked below). Daily climate projections simulated by the Canadian Regional Climate Model 4 (CanRCM4) forced by the second-generation Canadian Earth System Model (CanESM2) for two Representative Concentration Pathways (RCP4.5 or moderate emissions scenario and RCP8.5 or highest baseline emission scenario) are selected as inputs to the statistical downscaling models. The National Solar Radiation Database (NSRDB) is used to build and calibrate statistical models.

Array↗

Simulating Responses of Gravitational-Wave Instrumentation

Synthetic LISA is a computer program for simulating the responses of the instrumentation of the NASA/ESA Laser Interferometer Space Antenna (LISA) mission, the purpose of which is to detect and study gravitational waves. Synthetic LISA generates synthetic time series of the LISA fundamental noises, as filtered through all the time-delay-interferometry (TDI) observables. (TDI is a method of canceling phase noise in temporally varying unequal-arm interferometers.) Synthetic LISA provides a streamlined module to compute the TDI responses to gravitational waves, according to a full model of TDI (including the motion of the LISA array and the temporal and directional dependence of the arm lengths). Synthetic LISA is written in the C++ programming language as a modular package that accommodates the addition of code for specific gravitational wave sources or for new noise models. In addition, time series for waves and noises can be easily loaded from disk storage or electronic memory. The package includes a Python-language interface for easy, interactive steering and scripting. Through Python, Synthetic LISA can read and write data files in Flexible Image Transport System (FITS), which is a commonly used astronomical data format.

Armstrong, John↗

Meteor Shower Identification and Characterization with Python

The short development time associated with Python and the number of astronomical packages available have led to increased usage within NASA. The Meteoroid Environment Office in particular uses the Python language for a number of applications, including daily meteor shower activity reporting, searches for potential parent bodies of meteor showers, and short dynamical simulations. We present our development of a meteor shower identification code that identifies statistically significant groups of meteors on similar orbits. This code overcomes several challenging characteristics of meteor showers such as drastic differences in uncertainties between meteors and between the orbital elements of a single meteor, and the variation of shower characteristics such as duration with age or planetary perturbations. This code has been proven to successfully and quickly identify unusual meteor activity such as the 2014 kappa Cygnid outburst. We present our algorithm along with these successes and discuss our plans for further code development.

Moorhead, Althea↗

Open-Source Data Engineering at NASA: CCMC's Approach to Managing Petabyte-Scale Heliophysics Data

The Community Coordinated Modeling Center (CCMC) at NASA Goddard Space Flight Center (GSFC) leads heliophysics research by providing open access to numerous models and their outputs. Our resources are available on-demand and continuously updated with real-time data, covering sun-earth interactions across multiple domains. These domains include coronal, heliosphere, inner and global magnetosphere, ionosphere, thermosphere, and lower atmosphere interactions. Operating in a hybrid environment, CCMC utilizes both self-owned hardware and Amazon Web Services (AWS) cloud infrastructure. Managing petabytes of data across multiple locations necessitates robust data engineering solutions. To address this challenge, CCMC has adopted industry-standard and open-source tools. We use Apache Airflow as our primary data engineering platform, Python for scripting and data processing, and GitLab for version control and CI/CD. Additionally, we employ Kubernetes for containerized services, Grafana and Prometheus for metrics and monitoring, and Terraform and Puppet for reproducible infrastructure as code. This presentation will discuss lessons learned from our data engineering experiences, platforms evaluated but found unsuitable for our scientific data requirements, and specific techniques developed to enhance data transfer speed and reliability. By using these technologies effectively, CCMC continues to advance heliophysics research through efficient data management and open-access modeling.

space weather↗

Generating Models of the Flattop Critical Assembly for Benchmark Experiments with Python

Los Alamos National Laboratory has been performing nuclear criticality experiments since 1946 at the Pajarito site, starting the Los Alamos Critical Experiments Facility in 1948. A transition period occurred between 2004 and 2011 as operations moved to the National Criticality Experiments Research Center (NCERC), where criticality experiments are now performed. Criticality experiments are essential for determination and verification of nuclear data used in calculations and modeling—such as radiation transport codes—throughout the industry, enhancing nuclear criticality safety. In addition to nuclear data validation and benchmarking, the remotely operated critical assemblies at NCERC are used for a variety of experiments and training classes supporting criticality safety.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Truchas Overview

Truchas and Truchas-PBF are two sister codes for part-scale multi-physics modeling of manufacturing processes. Both programs are open source and made publicly available. They’re designed for efficient use of HPC resources and can be programmatically driven from Python packages. This enables automatic execution and analysis of ensembles of simulations, in some cases allowing 1000s of simulations to be evaluated in a day on HPC. Beyond just giving engineers a window into the concealed internal state of a system, the goal of Truchas is to provide a framework for developing novel manufacturing processes by understanding how the entire space of engineering inputs affects thermal state. It often is used to explore combinations of capabilities uncommon in commercial software, or to scale up analyses beyond the capabilities of commercial software.

97 MATHEMATICS AND COMPUTING↗

NCBI’s Virus Discovery Codeathon: Building “FIVE” —The Federated Index of Viral Experiments API Index

Viruses represent important test cases for data federation due to their genome size and the rapid increase in sequence data in publicly available databases. However, some consequences of previously decentralized (unfederated) data are lack of consensus or comparisons between feature annotations. Unifying or displaying alternative annotations should be a priority both for communities with robust entry representation and for nascent communities with burgeoning data sources. To this end, during this three-day continuation of the Virus Hunting Toolkit codeathon series (VHT-2), a new integrated and federated viral index was elaborated. This Federated Index of Viral Experiments (FIVE) integrates pre-existing and novel functional and taxonomy annotations and virus–host pairings. Variability in the context of viral genomic diversity is often overlooked in virus databases. As a proof-of-concept, FIVE was the first attempt to include viral genome variation for HIV, the most well-studied human pathogen, through viral genome diversity graphs. As per the publication of this manuscript, FIVE is the first implementation of a virus-specific federated index of such scope. FIVE is coded in BigQuery for optimal access of large quantities of data and is publicly accessible. Many projects of database or index federation fail to provide easier alternatives to access or query information. To this end, a Python API query system was developed to enhance the accessibility of FIVE.

59 BASIC BIOLOGICAL SCIENCES↗

pyRMG: A framework for high-throughput, large-cell DFT calculations on supercomputers

Exascale computing delivers the raw power to simulate ever larger and more chemically realistic systems, but realizing this potential requires codes that can efficiently use thousands of processors. Our real-space multigrid (RMG) density functional theory (DFT) code’s grid-decomposition approach scales nearly linearly with the number of graphics processing units (GPUs), even for simulations exceeding thousands of atoms. This scalability makes RMG a compelling tool for high-throughput DFT studies of materials that would otherwise be bottlenecked in other codes (for example, by global fast Fourier transforms in plane-wave DFT). However, the limited workflow infrastructure for RMG has thus far constrained its adoption to a small user community. In this work, we present pyRMG, a Python package designed to streamline the setup and execution of RMG DFT calculations. Built on the pymatgen and ASE (Atomic Simulation Environment) computational materials science Python packages, pyRMG automates input generation and convergence checking, and it integrates with modern job schedulers (e.g., Flux) on leadership-class platforms such as Frontier and Perlmutter. Here, we demonstrate pyRMG for a high-throughput study of strain effects in 2D 2L-Bi 2 Se 3 /2L-NbSe 2 heterostructures, which offers chemical insights into this system and shows that RMG-based workflows can converge with limited user intervention.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PyOED: An Extensible Suite for Data Assimilation and Model-Constrained Optimal Design of Experiments

This article describes PyOED, a highly extensible scientific package that enables developing and testing model-constrained optimal experimental design (OED) for inverse problems. Specifically, PyOED aims to be a comprehensive Python toolkit for model-constrained OED. The package targets scientists and researchers interested in understanding the details of OED formulations and approaches. It is also meant to enable researchers to experiment with standard and innovative OED technologies with a wide range of test problems (e.g., simulation models). OED, inverse problems (e.g., Bayesian inversion), and data assimilation (DA) are closely related research fields, and their formulations overlap significantly. Thus, PyOED is continuously being expanded with a plethora of Bayesian inversion, DA, and OED methods as well as new scientific simulation models, observation error models, and observation operators. These pieces are added such that they can be permuted to enable testing OED methods in various settings of varying complexities. The PyOED core is completely written in Python and utilizes the inherent object-oriented capabilities; however, the current version of PyOED is meant to be extensible rather than scalable. Specifically, PyOED is developed to “enable rapid development and benchmarking of OED methods with minimal coding effort and to maximize code reutilization.” This article provides a brief description of the PyOED layout and philosophy and provides a set of exemplary test cases and tutorials to demonstrate the potential of the package.

97 MATHEMATICS AND COMPUTING↗

ALchemist (Active Learning Toolkit for Chemical and Materials Research) [SWR-25-102]

ALchemist is a modular Python toolkit that brings active learning and Bayesian optimization to experimental design in chemical and materials research. It is designed for scientists and engineers who want to efficiently explore or optimize high-dimensional variable spaces—without writing code—using an intuitive graphical interface.

Coatney, Caleb [National Renewable Energy Laborato↗