Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Jupyter”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

PACE Water Resources: Demonstrating the Use of NASA's PACE Hyperspectral Ocean Color Instrument Data for Enhanced Coastal Management

This project developed tools to support the future use of Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) hyperspectral imagery in water resource monitoring and research by NASA DEVELOP teams and members of the PACE applications community. We sought to address a need for support in processing and visualizing hyperspectral PACE Ocean Color Instrument (OCI) data among researchers and decision-makers working in coastal water quality management and harmful algal bloom (HAB) monitoring. To supplement the day of simulated PACE imagery available, we used Aqua MODIS earth observations with Level 3 processing from March 2022 to build a Python graphical user interface (GUI) for visualizing ocean biogeochemical parameters relevant to the early detection and monitoring of HABs. We used simulated PACE OCI Level 2 data derived from the Python Top of Atmosphere Simulation Tool (PyTOAST) to build Jupyter Notebooks for band subset and selection. The Level 3 PACE Viewer components support users with quick visualizations as well as the creation of geoTIFFs and time-series. The Level 2 Jupyter Notebooks address users’ concerns over the volume and complexity of hyperspectral imagery. The PACE Viewer is useful for visual inspection and netCDF data processing but should not be used for geospatial analysis. Once PACE launches, this tool will alleviate the technical burdens of working with hyperspectral data and support the early detection and monitoring of HABs using PACE satellite imagery.

Python Top of Atmosphere Simulation Tool↗

CORAL: A framework for rigorous self-validated data modeling and integrative, reproducible data analysis

Abstract Background Many organizations face challenges in managing and analyzing data, especially when relevant datasets arise from multiple sources and methods. Analyzing heterogeneous datasets and additional derived data requires rigorous tracking of their interrelationships and provenance. This task has long been a Grand Challenge of data science and has more recently been formalized in the FAIR principles: that all data objects be Findable, Accessible, Interoperable, and Reusable, both for machines and for people. Adherence to these principles is necessary for proper stewardship of information, for testing regulatory compliance, for measuring the efficiency of processes, and for facilitating reuse of data-analytical frameworks. Findings We present the Contextual Ontology-based Repository Analysis Library (CORAL), a platform that greatly facilitates adherence to all 4 of the FAIR principles, including the especially difficult challenge of making heterogeneous datasets Interoperable and Reusable across all parts of a large, long-lasting organization. To achieve this, CORAL's data model requires that data generators extensively document the context for all data, and our tools maintain that context throughout the entire analysis pipeline. CORAL also features a web interface for data generators to upload and explore data, as well as a Jupyter notebook interface for data analysts, both backed by a common API. Conclusions CORAL enables organizations to build FAIR data types on the fly as they are needed, avoiding the expense of bespoke data modeling. CORAL provides a uniquely powerful platform to enable integrative cross-dataset analyses, generating deeper insights than are possible using traditional analysis tools.

97 MATHEMATICS AND COMPUTING↗

TomoPyUI : a user-friendly tool for rapid tomography alignment and reconstruction

The management and processing of synchrotron and neutron computed tomography data can be a complex, labor-intensive and unstructured process. Users devote substantial time to both manually processing their data ( i.e. organizing data/metadata, applying image filters etc. ) and waiting for the computation of iterative alignment and reconstruction algorithms to finish. In this work, we present a solution to these problems: TomoPyUI , a user interface for the well known tomography data processing package TomoPy . This highly visual Python software package guides the user through the tomography processing pipeline from data import, preprocessing, alignment and finally to 3D volume reconstruction. The TomoPyUI systematic intermediate data and metadata storage system improves organization, and the inspection and manipulation tools (built within the application) help to avoid interrupted workflows. Notably, TomoPyUI operates entirely within a Jupyter environment. Herein, we provide a summary of these key features of TomoPyUI , along with an overview of the tomography processing pipeline, a discussion of the landscape of existing tomography processing software and the purpose of TomoPyUI , and a demonstration of its capabilities for real tomography data collected at SSRL beamline 6-2c.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Open Chemistry, JupyterLab, REST, and quantum chemistry

Quantum chemistry must evolve if it wants to fully leverage the benefits of the internet age, where the worldwide web offers a vast tapestry of tools that enable users to communicate and interact with complex data at the speed and convenience of a button press. The Open Chemistry project has developed an open-source framework that offers an end-to-end solution for producing, sharing, and visualizing quantum chemical data interactively on the web using an array of modern tools and approaches. These tools build on some of the best open-source community projects such as Jupyter for interactive online notebooks, coupled with 3D accelerated visualization, state-of-the-art computational chemistry codes including NWChem and Psi4, and emerging machine learning and data mining tools such as ChemML and ANI. They offer flexible formats to import and export data, along with approaches to compare computational and experimental data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

BASIN-3D: A brokering framework to integrate diverse environmental data

Diverse observational and simulation datasets are needed to understand and predict complex ecosystem behavior over seasonal to decadal and century time-scales. Integration of these datasets poses a major barrier towards advancing environmental science, particularly due to differences in the structure and formats of data provided by various sources. Here, we describe BASIN-3D (Broker for Assimilation, Synthesis and Integration of eNvironmental Diverse, Distributed Datasets), a data integration framework designed to dynamically retrieve and transform heterogeneous data from different sources into a common format to provide an integrated view. BASIN-3D enables users to adopt a standardized approach for data retrieval and avoid customizations for the data type or source. We demonstrate the value of BASIN-3D with two use cases that require integration of data from regional to watershed spatial scales. The first application uses the BASIN-3D Python library to integrate time-series hydrological and meteorological data to provide standardized inputs to analytical and machine learning codes in order to predict the impacts of hydrological disturbances on large river corridors of the United States. The second application uses the BASIN-3D Django framework to integrate diverse time-series data in a mountainous watershed in East River, Colorado, United States to enable scientific researchers to explore and download data through an interactive web portal. Thus, BASIN-3D can be used to support data integration for both web-based tools, as well as data analytics using Python scripting and extensions like Jupyter notebooks. The framework is expected to be transferable to and useful for many other field and modeling studies.

Varadharajan, C↗

Graph neural networks for CO 2 solubility predictions in Deep Eutectic Solvents

Deep Eutectic Solvents (DESs) are a promising class of solvents for CO 2 capture. DESs are complex mixtures that can be designed to optimize CO solubility and overall capture process efficiency. However, the vast design landscape of DES mixtures makes experimental investigation prohibitive; as such, there is a need for computational models that can quickly and efficiently navigate the design space and inform data collection efforts. In this work, we propose Graph Neural Network (GNN) models for predicting CO 2 solubility for DESs; the GNN leverages a mixture graph representation that captures the molecular structure of the DES components as well as their intermolecular interactions. Here, we compare the GNN framework against alternative architectures (neural networks, graph convolution networks, and random forests) and data representations (molecular fingerprints, sigma profiles, and graphs). We show that the proposed approach offers superior predictive performance; specifically, we show that solubility can be predicted reliably directly from molecular structure (without the need of using sigma profiles as proposed in previous studies). This result is important, as obtaining sigma profiles requires expensive density functional theory computations. We also explored the ability of GNNs to predict solubility for new DES mixtures and operating conditions. We found that the model extrapolates across temperature reliably. However, we also found deficiencies in the ability of the model to predict solubility for DES mixtures, pressures, and molar ratio not included in the training sets; we show that this is due to an inherent lack of chemical diversity in datasets available in the literature. The proposed computational capabilities can thus help navigate the design space of DES and inform data collection efforts. Our models, data, and benchmarks are shared as Python code implemented in Jupyter notebooks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The Baghdad Atlas: A relational database of inelastic neutron-scattering (n,n ' γ) data

A relational database has been developed based on the original (n,n'γ) work carried out by A. M. Demidov et al., at the Nuclear Research Institute in Baghdad, Iraq (Demidov et al., 1978) for 105 independent measurements comprising 76 elemental samples of natural composition and 29 isotopically-enriched samples. The information from this Atlas includes: γ-ray energies and relative intensities; nuclide and level data corresponding to the residual nucleus and meta data associated with the target sample that allows for the extraction of the flux-weighted (n,n'γ) cross sections for a given transition relative to a defined value. The optimized angular-distribution-corrected fast-neutron flux-weighted partial γ-ray cross section for the production of the 846.8-keV 21+→0gs+γ-ray transition in 56Fe, determined to be $\langle$σγ$\rangle$=143(29) mb, is used for this purpose. However, different values for the adopted cross section can be readily implemented to accommodate user preference based on revised determinations of this quantity. The Atlas (n,n'γ) data has been compiled into a series of CSV-style ASCII data sets and a suite of Python scripts have been developed to build and install the database locally. The database can then be accessed directly through the SQLite engine, or using alternative methods such as the Jupyter Notebook Python-browser interface. Several examples exploiting different interaction methodologies are distributed with the complete software package.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A methodology for decay heat characterization in molten salt reactors

Accurate decay heat prediction in molten salt reactors (MSRs) faces dual challenges: complex operational uncertainties and the need for interpretable models compatible with engineering workflows. This work presents a hybrid machine learning and segmented polynomial methodology that addresses both requirements through three key innovations. First, a modular data architecture encodes MSR-specific operational parameters (power density: 1-100 W cm -3 , humidity: 0-0.1 wt %, air ingress: 0-0.1 mol %) with uncertainty-aware temporal discretization spanning 15 orders of magnitude. Second, region-optimized machine learning models achieve 92.3 % root mean square error (RMSE) reduction over conventional polynomials while maintaining physical interpretability through automated piecewise equation generation. Third, dual front-end interfaces accelerate safety analyses — a Jupyter environment enables researchers to explore 10,000+ parameter combinations via interactive widgets, while a Streamlit web application reduces design iteration cycles through production-grade visualization tools. Operational deployment demonstrates prediction times of only a couple hundred milliseconds for 10 4 years decay profiles, enabling real-time optimization of spent fuel container designs.

42 - ENGINEERING↗

A FAIR-compliant parts catalogue for genome engineering and expression control in Saccharomyces cerevisiae

The synthetic biology toolkit for baker's yeast, Saccharomyces cerevisiae, includes extensive genome engineering toolkits and parts repositories. However, with the increasing complexity of engineering tasks and versatile applications of this model eukaryote, there is a continued interest to expand and diversify the rational engineering capabilities in this chassis by FAIR (findable, accessible, interoperable, and reproducible) compliance. In this study, we designed and characterised 41 synthetic guide RNA sequences to expand the CRISPR-based genome engineering capabilities for easy and efficient replacement of genomically encoded elements. Moreover, we characterize in high temporal resolution 20 native promoters and 18 terminators using fluorescein and LUDOX CL-X as references for GFP expression and OD600 measurements, respectively. Additionally, all data and reported analysis is provided in a publicly accessible jupyter notebook providing a tool for researchers with low-coding skills to further explore the generated data as well as a template for researchers to write their own scripts. We expect the data, parts, and databases associated with this study to support a FAIR-compliant resource for further advancing the engineering of yeasts.

59 BASIC BIOLOGICAL SCIENCES↗

Three-dimensional paganica fault morphology obtained from hypocenter clustering (L'Aquila 2009 seismic sequence, Central Italy)

In seismic modelling, fault planes are normally assumed to be flat due to the lack of data which can constrain fault morphology. However, incorporating 3D fault morphology is important for modelling several phenomena, for example calculating mainshock induced stress changes. Here we utilize a data-analytical method to unveil the 3D rupture morphology of faults using unsupervised clustering techniques applied to earthquake hypocenters in seismic sequences. We apply this method to the 2009 L'Aquila seismic sequence which involved a M W 6.1 mainshock on April 6th. We use a dataset of about 50,000 relocated events, mostly microearthquakes, reaching magnitude of completeness equal to 0.7. Clustering distinguishes the earthquakes as occurring in three main clusters along with other minor fault segments. We then represent the morphology of the main Paganica fault system (responsible for the largest mainshock) using splines. This method shows promise as a step toward robustly and quickly obtaining 3D rupture morphologies where earthquake sequences have been monitored. The 3D model is presented interactively online, and the processing is presented in an interactive Jupyter Notebook (https://bit.ly/2MnCFdj).

58 GEOSCIENCES↗

Identification and correction of temporal and spatial distortions in scanning transmission electron microscopy

Scanning transmission electron microscopy (STEM) has become the technique of choice for quantitative characterization of atomic structure of materials, where the minute displacements of atomic columns from high-symmetry positions can be used to map strain, polarization, octahedra tilts, and other physical and chemical order parameter fields. The latter can be used as inputs into mesoscopic and atomistic models, providing insight into the correlative relationships and generative physics of materials on the atomic level. However, these quantitative applications of STEM necessitate understanding the microscope induced image distortions and developing the pathways to compensate them both as part of a rapid calibration procedure for in situ imaging, and the post-experimental data analysis stage. Here, we explore the spatiotemporal structure of the microscopic distortions in STEM using multivariate analysis of the atomic trajectories in the image stacks. Based on the behavior of principal component analysis (PCA), we develop the Gaussian process (GP)-based regression method for quantification of the distortion function. The limitations of such an approach and possible strategies for implementation as a part of in-line data acquisition in STEM are discussed. Here, the analysis workflow is summarized in a Jupyter notebook that can be used to retrace the analysis and analyze the reader's data.

36 MATERIALS SCIENCE↗

Cluster Analysis of Combined EDS and EBSD Data to Solve Ambiguous Phase Identifications

A common problem in analytical scanning electron microscopy (SEM) using electron backscatter diffraction (EBSD) is the differentiation of phases with distinct chemistry but the same or very similar crystal structure. X-ray energy dispersive spectroscopy (EDS) is useful to help differentiate these phases of similar crystal structures but different elemental makeups. However, open, automated, and unbiased methods of differentiating phases of similar EBSD responses based on their EDS response are lacking. This paper describes a simple data analytics-based method, using a combination of singular value decomposition and cluster analysis, to merge simultaneously acquired EDS + EBSD information and automatically determine phases from both their crystal and elemental data. I use hexagonal TiB 2 ceramic contaminated with multiple crystallographically ambiguous but chemically distinct cubic phases to illustrate the method. Code, in the form of a Python 3 Jupyter Notebook, and the necessary data to replicate the analysis are provided as Supplementary material.

47 OTHER INSTRUMENTATION↗

Machine Learning for Materials Scientists: An Introductory Guide toward Best Practices

This Methods/Protocols article is intended for materials scientists interested in performing machine learning-centered research. Herein, we cover broad guidelines and best practices regarding the obtaining and treatment of data, feature engineering, model training, validation, evaluation and comparison, popular repositories for materials data and benchmarking data sets, model and architecture sharing, and finally publication. In addition, we include interactive Jupyter notebooks with example Python code to demonstrate some of the concepts, workflows, and best practices discussed. Overall, the data-driven methods and machine learning workflows and considerations are presented in a simple way, allowing interested readers to more intelligently guide their machine learning research using the suggested references, best practices, and their own materials domain expertise.

36 MATERIALS SCIENCE↗

A FAIR and AI-ready Higgs boson decay dataset

Abstract To enable the reusability of massive scientific datasets by humans and machines, researchers aim to adhere to the principles of findability, accessibility, interoperability, and reusability (FAIR) for data and artificial intelligence (AI) models. This article provides a domain-agnostic, step-by-step assessment guide to evaluate whether or not a given dataset meets these principles. We demonstrate how to use this guide to evaluate the FAIRness of an open simulated dataset produced by the CMS Collaboration at the CERN Large Hadron Collider. This dataset consists of Higgs boson decays and quark and gluon background, and is available through the CERN Open Data Portal. We use additional available tools to assess the FAIRness of this dataset, and incorporate feedback from members of the FAIR community to validate our results. This article is accompanied by a Jupyter notebook to visualize and explore this dataset. This study marks the first in a planned series of articles that will guide scientists in the creation of FAIR AI models and datasets in high energy particle physics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Implementation and analysis of quantum computing application to Higgs boson reconstruction at the large Hadron Collider

With the advent of the High-Luminosity Large Hadron Collider (HL-LHC) era, high energy physics (HEP) event selection will require new approaches to rapidly and accurately analyze vast databases. The current study addresses the enormity of HEP databases in an unprecedented manner—a quantum search using Grover’s Algorithm (GA) on an unsorted database, ATLAS Open Data, from the ATLAS detector. A novel method to identify rare events at 13 TeV in CERN’s LHC using quantum computing (QC) is presented. As indicated by the Higgs boson decay channel H→ZZ*→4l , the detection of four leptons in one event may be used to reconstruct the Higgs boson and, more importantly, evince Higgs boson decay to some new phenomena, such as H→ZZ d →4l . Searching the dataset for collisions resulting in detection of four leptons using a Jupyter Notebook, a classical simulation of GA, and several quantum computers with multiple qubits, the current application was found to make the proper selection in the unsorted dataset. Quantum search efficacy was analyzed for the incoming HL-LHC by implementing the QC method on multiple classical simulators and IBM’s quantum computers with the IBM Qiskit Open Source Software. The current QC application provides a novel, high-efficiency alternative to classical database searches, demonstrating its potential utility as a rapid and increasingly accurate search method in HEP.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Discovering invariant spatial features in electron energy loss spectroscopy images on the mesoscopic and atomic levels

Over the last two decades, Electron Energy Loss Spectroscopy (EELS) imaging with a scanning transmission electron microscope has emerged as a technique of choice for visualizing complex chemical, electronic, plasmonic, and phononic phenomena in complex materials and structures. The availability of the EELS data necessitates the development of methods to analyze multidimensional data sets with complex spatial and energy structures. Traditionally, the analysis of these data sets has been based on analysis of individual spectra, one at a time, whereas the spatial structure and correlations between individual spatial pixels containing the relevant information of the physics of underpinning processes have generally been ignored and analyzed only via the visualization as 2D maps. Here, we develop a machine learning-based approach and workflows for the analysis of spatial structures in 3D EELS data sets using a combination of dimensionality reduction and multichannel rotationally invariant variational autoencoders. This approach is illustrated for the analysis of both the plasmonic phenomena in a system of nanowires and in the core excitations in functional oxides using low loss and core-loss EELS, respectively. The code developed in this manuscript is open sourced and freely available and provided as a Jupyter notebook for the interested reader.

36 MATERIALS SCIENCE↗

PigmentHunter: A point-and-click application for automated chlorophyll-protein simulations

Chlorophyll proteins (CPs) are the workhorses of biological photosynthesis, working together to absorb solar energy, transfer it to chemically active reaction centers, and control the charge-separation process that drives its storage as chemical energy. Yet predicting CP optical and electronic properties remains a serious challenge, driven by the computational difficulty of treating large, electronically coupled molecular pigments embedded in a dynamically structured protein environment. To address this challenge, we introduce here an analysis tool called PigmentHunter, which automates the process of preparing CP structures for molecular dynamics (MD), running short MD simulations on the nanoHUB.org science gateway, and then using electrostatic and steric analysis routines to predict optical absorption, fluorescence, and circular dichroism spectra within a Frenkel exciton model. Inter-pigment couplings are evaluated using point-dipole or transition-charge coupling models, while site energies can be estimated using both electrostatic and ring-deformation approaches. The package is built in a Jupyter Notebook environment, with a point-and-click interface that can be used either to manually prepare individual structures or to batch-process many structures at once. Here, we illustrate PigmentHunter’s capabilities with example simulations on spectral line shapes in the light harvesting 2 complex, site energies in the Fenna–Matthews–Olson protein, and ring deformation in photosystems I and II.

14 SOLAR ENERGY↗

Veridical data science

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, composed of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. As part of the PCS workflow, we develop PCS inference procedures, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others. Moreover, we demonstrate its favorable performance over existing methods in terms of receiver operating characteristic (ROC) curves in high-dimensional, sparse linear model simulations, including a wide range of misspecified models. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.

97 MATHEMATICS AND COMPUTING↗