Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open source tools”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Reverse engineering environmental metatranscriptomes clarifies best practices for eukaryotic assembly

Abstract Background Diverse communities of microbial eukaryotes in the global ocean provide a variety of essential ecosystem services, from primary production and carbon flow through trophic transfer to cooperation via symbioses. Increasingly, these communities are being understood through the lens of omics tools, which enable high-throughput processing of diverse communities. Metatranscriptomics offers an understanding of near real-time gene expression in microbial eukaryotic communities, providing a window into community metabolic activity. Results Here we present a workflow for eukaryotic metatranscriptome assembly, and validate the ability of the pipeline to recapitulate real and manufactured eukaryotic community-level expression data. We also include an open-source tool for simulating environmental metatranscriptomes for testing and validation purposes. We reanalyze previously published metatranscriptomic datasets using our metatranscriptome analysis approach. Conclusion We determined that a multi-assembler approach improves eukaryotic metatranscriptome assembly based on recapitulated taxonomic and functional annotations from an in-silico mock community. The systematic validation of metatranscriptome assembly and annotation methods provided here is a necessary step to assess the fidelity of our community composition measurements and functional content assignments from eukaryotic metatranscriptomes.

Krinos, Arianna I. (ORCID:0000000197678392)↗

The "PVLib" of Degradation: PVDeg

The Photovoltaic (PV) industry constantly aims for lower costs through higher-efficiency cells, improved module designs, and improvements in durability. This leads to the use of new materials, designs, and manufacturing processes, and not always with a sufficient amount of durability testing. To help drive down costs there is a desire to create modules that will last for up to 50 years of service life. To accomplish this, every degradation mode and mechanism must be identified and either eliminated or otherwise mitigated. This involves the extrapolation of laboratory results to the field conditions. There is a need to organize the existing degradation data into an accessible format and to provide industry relevant tools for extrapolation from laboratory to field conditions. While the basic equations used to model degradation are sometimes very simple, the full analysis involves calculations are cumbersome but ubiquitous for many degradation processes. A simplified, modeling framework to accomplish these repetitive processes will facilitate the analysis to help researchers keep up with the rapid pace of technological changes. In this talk, we will describe our progress creating the open-source tool PVDeg. This tool can be used to search for and analyze degradation information and extrapolate PV module performance and durability to field exposure. PVDeg simplifies many of the common foundational computational operations for obtaining meteorological data and using it to generate a model of the PV deployment. This prediction tool repository also contains various degradation models as well as a library of material parameters suitable for estimating the durability assessment of materials and components. We use an integration pipeline approach that allows us to leverage weather data from the National Solar Radiation Database, and other weather sources, to perform geospatial degradation analysis in the US and worldwide. We hope to become a repository that can be used for weathering and degradation analysis for various applications beyond the PV industry. During the talk, we will provide the PVPMC attendees the opportunity to interact with the tool via a Google Collab tutorial they can run on their phones or laptops.

durability↗

Adiabatic Gravitational Waveform Model for Compact Objects Undergoing Quasi-Circular Inspirals Into Rotating Massive Black Holes

We present bhpwave: a new python-based, open-source tool for generating the gravitational waveforms of stellar-mass compact objects undergoing quasicircular inspirals into rotating massive black holes. These binaries, known as extreme-mass-ratio inspirals (EMRIs), are exciting mHz gravitational wave sources for future space-based detectors such as the Laser Interferometer Space Antenna (LISA). Relativistic models of EMRI gravitational wave signals are necessary to unlock the full scientific potential of mHz detectors, yet few open-source EMRI waveform models exist. Thus we built bhpwave, which uses the adiabatic approximation from black hole perturbation theory to rapidly construct gravitational waveforms based on the leading-order inspiral dynamics of the binary. In this work, we present the theoretical and numerical foundations underpinning bhpwave. We also demonstrate how bhpwave can be used to assess the impact of EMRI modeling errors on LISA gravitational wave data analysis. In particular, we find that for retrograde orbits and slowly spinning black holes we can mismodel the gravitational wave phasing by as much as ∼10 radians without significantly biasing EMRI parameter estimation.

Zachary Nasipak↗

Implementation and evaluation of multi-dual mode counter-current chromatography in the CUP Modeler software

Counter-current chromatography (CCC) is a separation technique that utilizes immiscible solvent pairs as stationary and mobile phases, which imparts numerous benefits compared to solid-liquid chromatography including the ability to treat either the more-dense or less-dense solvent layer as the mobile phase. Multi-dual mode (MDM) is a CCC elution mode capable of improving the separation of closely eluting compounds by alternating upper- and lower-layer solvent flows in opposing directions within the same separation. While some effort has been made to model MDM, implementation of these models in experimental design has yet to be widely adopted. Accordingly, we further developed our previously published cell utilized partitioning (CUP) model to include MDM predictions with CCC and packaged the full suite of CUP modeling capabilities into a user-friendly, open-source tool called the CUP Modeler. The mathematical model for MDM CCC was derived and validated with experimental separation of ethyl guaiacol (EG) and ethyl phenol (EP), two compounds that co-elute in our previously demonstrated reductive catalytic fractionation (RCF) lignin monomer isolation method. The developed MDM model provided insights into the effect of multiple operating parameters - including stationary phase retention, flow rate, column efficiency, feed concentration ratio, selectivity factor, and solute distribution ratios - on the separation yields, productivity, and purities. Our model agreed with prevailing understanding of MDM but also revealed new insights including that the ideal distribution ratios for co-eluting solutes to be separated by MDM is between 1.1 and 1.5, with the lower value ideally close to 1.25. Overall, this work provides fundamental insights for MDM process design and enables broader adoption of general liquid-liquid chromatography with a new, open-source user-friendly interface.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Creating a Simulation Platform for Research and Development of Advanced Control Methods

Advanced nuclear reactors are essential to meet the changing energy requirements throughout both the United States and the rest of world. In addition to other features, they are designed to enable deployment in remote locations and operate in a fully (or near-fully) autonomous manner, which will require a new control paradigm. To realize autonomously operating reactors, the U.S. Department of Energy’s Nuclear Energy Enabling Technologies Advanced Sensors and Instrumentation (NEET ASI) program conducts research and development into the enabling technologies and methods needed, including digital twins, machine learning, and risk modeling, in addition to various types of control methods. These technologies and methods are the key foundations needed to achieve fully autonomous systems. To develop and evaluate the technologies and methods necessary for achieving autonomous operations, it is critical to identify a software tool capable of integrating all the required elements. In surveying the available solutions, no software platforms were identified that could accomplish what was needed without introducing drawbacks. This challenge was the motivation for the current effort: to develop a software platform that can seamlessly integrate autonomouscontrol-enabling technologies and methods, allowing for accelerated research and development and transfer of ideas. The resulting platform, known as the Control and Optimization Modular Modeling Application for Nuclear Deployment (COMMAND), is Python-based, and leverages open-source tools to provide flexibility and facilitate building upon prior research. It is designed to enable advanced reactor developers to deploy and test advanced control technologies and methods coupled with their own models, solutions, and hardware. Given the substantial undertaking of developing such a platform, the current effort focused on laying down scalable, flexible software foundations and infrastructure, then demonstrating the platform via a use case. These foundations included developing generic modules, which contain the base variable and system blocks (the information and functional building blocks, respectively, that can be used to design a simulation) and the data handling and storage blocks needed to exchange information between the various blocks; as well as enablingtechnology-specific modules. This platform was evaluated via a use case, which was to simulate and control a process for the Microreactor Automated Control System (MACS) test bed. While MACS is not currently directly coupled to any specific microreactor physics, it was initially developed in concert with the Microreactor Applications Research Validation and Evaluation (MARVEL) microreactor, and so the MARVEL physics are used here. As part of this use case, several enabling-technology-specific blocks within COMMAND were integrated, including a proportional integral derivative (PID) control block, a Reactor Excursion and Leak Analysis Program (RELAP5-3D) block, and an anomaly detection block. The COMMAND software platform was successfully demonstrated to achieve the scalability and flexibility objectives of this effort and will be leveraged by the program’s research efforts to advance state of the art control methodologies towards autonomous operations of advanced reactors. As new use cases are created and implemented, it is anticipated that COMMAND will continue to grow and evolve to meet new requirements.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

MLOps for Beam Controls

Machine learning operations (MLOps) is the standardization and streamlining of the ML development lifecycle to address the challenges associated with large-scale machine learning applications. The full MLOps pipeline consists of open-source tools: DataHub, MinIO and MLflow. It is being used for dataset management and model development to handle changing data dependencies, varying business needs, reproducibility, and diverse teams working with differing tools and skills. To demonstrate the completion of an MLOps pipeline for particle accelerator operations, we are deploying a simple script that computes settings for the Booster’s gradient magnet power supply. Once the demonstration is complete, we will develop and deploy ML-based optimization algorithms to improve Booster’s overall efficiency. This MLOps pipeline opens the gate to systematically develop and deploy ML applications for accelerator controls and diagnostics.

43 PARTICLE ACCELERATORS↗

MLOps for Beam Controls

Machine learning operations (MLOps) is the standardization and streamlining of the ML development lifecycle to address the challenges associated with large-scale machine learning applications. The full MLOps pipeline consists of open-source tools: DataHub, MinIO and MLflow. It is being used for dataset management and model development to handle changing data dependencies, varying business needs, reproducibility, and diverse teams working with differing tools and skills. To demonstrate the completion of an MLOps pipeline for particle accelerator operations, we are deploying a simple script that computes settings for the Booster’s gradient magnet power supply. Once the demonstration is complete, we will develop and deploy ML-based optimization algorithms to improve Booster’s overall efficiency. This MLOps pipeline opens the gate to systematically develop and deploy ML applications for accelerator controls and diagnostics.

43 PARTICLE ACCELERATORS↗

MLOps for Beam Controls

Machine learning operations (MLOps) is the standardization and streamlining of the ML development lifecycle to address the challenges associated with large-scale machine learning applications. The full MLOps pipeline consists of open-source tools: DataHub, MinIO and MLflow. It is being used for dataset management and model development to handle changing data dependencies, varying business needs, reproducibility, and diverse teams working with differing tools and skills. To demonstrate the completion of an MLOps pipeline for particle accelerator operations, we are deploying a simple script that computes settings for the Booster’s gradient magnet power supply. Once the demonstration is complete, we will develop and deploy ML-based optimization algorithms to improve Booster’s overall efficiency. This MLOps pipeline opens the gate to systematically develop and deploy ML applications for accelerator controls and diagnostics.

43 PARTICLE ACCELERATORS↗

RAVIS: Resource Forecast and Ramp Visualization for Situational Awareness of Variable Renewable Generation

The Resource Forecast and Ramp Visualization for Situational Awareness (RAVIS) is an innovative, open-source tool for visualizing variable renewable resource forecasts and alerts for significant up- and down-ramps in renewable generation and the consequent system net load. The modular dashboard of RAVIS contains configurable panes for viewing probabilistic time-series forecasts, ramp event alerts on the look-ahead timeline, and spatially resolved resource sites, and forecasts. For comprehensive situational awareness, the tool can add additional data layers from simulation and independent system operator (ISO) market clearing data---including electric line utilization, nodal price, and available generation flexibility---in response to continuously updated renewable forecasts. This paper introduces the RAVIS technology suite employed to provide optimum visualization and flexible design characteristics. The paper also illustrates some use cases of the tool using site-specific solar power forecast data obtained from the IBM Watt-Sun forecasting platform for the California ISO and Midcontinent ISO footprints. The source code for RAVIS is public (https://github.com/ravis-nrel/ravis), and the intended users are forecasters, utility planners, ISO operators, and researchers.

41 EE - Solar Energy Technologies Office (EE-4S)↗

RAVIS: Resource Forecast and Ramp Visualization for Situational Awareness of Variable Renewable Generation: Preprint

The Resource Forecast and Ramp Visualization for Situational Awareness (RAVIS) is an innovative, open-source tool for visualizing variable renewable resource forecasts and alerts for significant up- and down-ramps in renewable generation and the consequent system net load. The modular dashboard of RAVIS contains configurable panes for viewing probabilistic time-series forecasts, ramp event alerts on the look-ahead timeline, and spatially resolved resource sites, and forecasts. For comprehensive situational awareness, the tool can add additional data layers from simulation and independent system operator (ISO) market clearing data—including electric line utilization, nodal price, and available generation flexibility—in response to continuously updated renewable forecasts. This paper introduces the RAVIS technology suite employed to provide optimum visualization and flexible design characteristics. The paper also illustrates some use cases of the tool using site-specific solar power forecast data obtained from the IBM Watt-Sun forecasting platform for the California ISO and Midcontinent ISO footprints. The source code for RAVIS is public (https://github.com/ravis-nrel/ravis), and the intended users are forecasters, utility planners, ISO operators, and researchers.

41 EE - Solar Energy Technologies Office (EE-4S)↗

A Bottom-Up Cost Estimation Tool for Nuclear Microreactors

The rising interest in nuclear microreactors has highlighted the need for comprehensive technoeconomic assessments. However, the scarcity of publicly available designs and cost data has posed significant challenges. To address this issue, the Microreactor Optimization Using Simulation and Economics (MOUSE) tool is developed. MOUSE is a tool that integrates nuclear microreactor design with reactor economics. The design calculations encompass core simulations using the OpenMC Monte Carlo Particle Transport Code [romano2015], along with simplified balance of plant calculations. On the economic side, MOUSE provides detailed bottom-up cost estimates, calculating both the total capital cost and the levelized cost of energy for first-of-a-kind and nth-of-a-kind microreactors. The cost estimation correlations are developed using data from the MARVEL project and additional literature sources. MOUSE has released as an open-source tool on GitHub (MOUSE Tool). By combining design calculations with cost equations, MOUSE enables users to evaluate the impact of various technological consideration, advanced moderators, design changes, material/fuel changes, and geometry modifications—as well as economic parameters like interest rates and construction duration. This comprehensive framework can guide stakeholders towards technological solutions that enhance microreactor competitiveness. Additionally, powered by the WATTS toolkit [romano2022], MOUSE supports optimization studies, parametric analyses, and uncertainty calculations/propagation. Currently, preconceptual designs of three microreactor types are included in MOUSE: a liquid metal thermal microreactor (LTMR), gas cooled TRISO-fueled microreactor (GCMR) and heat-pipe TRISO fueled microreactor (HPMR). To showcase its ability, MOUSE was used to conduct detailed bottom-up cost estimates for the first of a kind (FOAK) and Nth of a kind (NOAK) of the following microreactors • A 20MWt LTMR that is built on the ongoing MARVEL demonstration at Idaho National Laboratory (INL) • A 15 MWt GCMR that was designed to be more representative of the typical commercial microreactor • A 7 MWt HPMR that was built on previous work (Choi 2024) The The reader should note that these three designs and corresponding cost estimates are examples to demonstrate the MOUSE capability. The designs are pre-conceptual, the reactor designs were not optimized, and the cost estimates were developed with incomplete information. Additionally, stakeholders might be interested in a variety of designs that may differ from the examples provided in this report. The MOUSE tool can also be used to study how design choices affect economics. To demonstrate its capability, MOUSE was used conduct parametric studies such as examining the economic impact of the reflector's material and thickness, the moderator's booster material and dimensions, fuel composition and enrichment, core size, and power level. Several insights were gained from these parametric studies.

Hanna, Botros↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Open-Source Data Engineering at NASA: CCMC's Approach to Managing Petabyte-Scale Heliophysics Data

The Community Coordinated Modeling Center (CCMC) at NASA Goddard Space Flight Center (GSFC) leads heliophysics research by providing open access to numerous models and their outputs. Our resources are available on-demand and continuously updated with real-time data, covering sun-earth interactions across multiple domains. These domains include coronal, heliosphere, inner and global magnetosphere, ionosphere, thermosphere, and lower atmosphere interactions. Operating in a hybrid environment, CCMC utilizes both self-owned hardware and Amazon Web Services (AWS) cloud infrastructure. Managing petabytes of data across multiple locations necessitates robust data engineering solutions. To address this challenge, CCMC has adopted industry-standard and open-source tools. We use Apache Airflow as our primary data engineering platform, Python for scripting and data processing, and GitLab for version control and CI/CD. Additionally, we employ Kubernetes for containerized services, Grafana and Prometheus for metrics and monitoring, and Terraform and Puppet for reproducible infrastructure as code. This presentation will discuss lessons learned from our data engineering experiences, platforms evaluated but found unsuitable for our scientific data requirements, and specific techniques developed to enhance data transfer speed and reliability. By using these technologies effectively, CCMC continues to advance heliophysics research through efficient data management and open-access modeling.

space weather↗

An Open-Source Framework for PV in the Circular Economy Evaluation

This open-source tool follows the dynamic Material and flows for each component and material in a PV system, from mining to end of life to evaluate the material and energy impacts of different paths. Following circularity principles, virgin material inputs can be offset by recovered materials to reduce impacts due to mining and extraction. Unlike consumer products, renewable energy technologies generate power over their useful life to quickly offset energy required for manufacturing. The model provides unique baselines of PV evolution, and allows us to evaluate the material and soon energy return on investment of decisions such as field repair, off-site refurbishment, reuse, or recycling. Energy layer and cost, as well as visual dashboard under development.

circular economy modeling↗

How Certain Physical Considerations Impact Aerostructural Wing Optimization

Wing design optimization has been studied extensively and is of continued interest as optimization tools are developed and become more accessible. In each of these studies, certain assumptions and simplifications are made to make the design problem tractable. However, it is difficult to find systematic studies in which several considerations are added or removed one at a time to study how much impact they have. In this work, we examine how certain physical considerations (viscous drag, wave drag, thrust loads, and inertial relief from structural, fuel, and engine masses), impact the aerostructural optimization results for three distinct aircraft wings. The goal is to help develop a rough idea of how important these physical considerations are. We do this using gradient-based optimization and a multidisciplinary design optimization framework, OpenMDAO. We use the open-source tool OpenAeroStruct that couples a vortex lattice method to a finite element method. We establish a baseline aerostructural design optimization problem then perform a series of optimizations, each with one physical consideration removed from the baseline case. We find that depending on the size of the aircraft and flight conditions, the importance of some of these physical considerations varies considerably whereas the importance of others do not. Specifically, the optimal designs change radically without proper viscous and wave drag considerations and smaller aircraft with more distributed propulsion are more affected by the inclusion of engine loads.

Aeropropulsion↗

Automated Energy-Dispersive X-ray Spectroscopy Analysis for Multi-Modal Few-Shot Learning

Scanning transmission electron microscopy (STEM) is a powerful tool that allows for the atomic-scale analysis of a materials’ structure, chemistry, and defect domains (Akers et al. 2021). The current generation of microscopes generate vast amounts of data, surpassing the limits of effective manual analysis traditionally performed by domain experts (Spurgeon et al. 2021). While recent strides in machine learning have significantly enhanced the processing of large and intricate datasets acquired through electron microscopy, the prevalent use of proprietary software packages for initial data collection poses a challenge. In many cases, these software packages act as a ‘black box’, constraining user functionality and hindering the output of data in a format that is conducive to seamless integration into machine learning models. This work addresses these challenges by adapting HyperSpy, an open-source Python library, for the analysis and quantification of raw energy dispersive spectroscopy (EDS) data acquired through STEM. The modified HyperSpy code successfully facilitates user-defined segmentation of the data, enabling the integration of atomic %, weight %, and raw EDS spectra for each segmented region into an existing few-shot machine learning model. While initial results reveal discrepancies in quantified atomic and weight percentages when compared to proprietary software, ongoing efforts aim to rectify this issue by refining the fit of the HyperSpy model to the EDS spectra. Overall, this research underscores the potential of open-source tools like HyperSpy to enhance the accessibility of analytical tools, fostering a transparent and user-friendly environment for seamlessly incorporating electron microscopy data into machine learning models.

36 MATERIALS SCIENCE↗

INCREASING THE TRANSPARENCY AND REPRODUCIBILITY OF SPACE RADIATION SCIENCE: THE RADIATION BIOLOGY ONTOLOGY

Among the primary objectives of the Open/Open-Source Science paradigm are making scientific investigation data transparent and results reproducible [1], objectives shared by the FAIR principles [2]. To accomplish this, the conceptual framework that includes all the investigation objects needs to be accurately captured and communicated to all data consumers. A large part of this requires using metadata standards to annotate data collected. These standards should be readily accessible, informed by scientific community consensus and sufficiently specific to encompass all of the important aspects of the investigation. Starting in 2020 we have been co-leading an open consortium to develop a new metadata standard, the Radiation Biology Ontology (RBO), through the Open Biological and Biomedical Ontologies (OBO) Foundry [3]. We began by transforming many of the terms from the National Council on Radiation Protection and Measurement into concepts that can be formally related to existing OBO Foundry classes or attributes. We then identified and imported into the RBO existing OBO Foundry classes that have obvious relevance for radiation biomedicine (for example, concepts from the Environment Ontology that describe radiative processes, and concepts from the Gene Ontology dealing with molecular and cellular responses to radiation). Finally, we scrutinized datasets from investigations of radiation effects held in NASA GeneLab and LSDA repositories and added additional classes, instances, and attributes into the RBO that should be used to annotate these data. We developed the RBO using the open-source tools of GitHub and publish the RBO periodically through the NIH/NCBI BioPortal website, so systems worldwide can leverage the knowledge it contains [4]. This initial phase of concept modeling has yielded an RBO that at present has more than 300 declared concepts, with more than 3500 additional concepts imported from other OBO Foundry ontologies. While this first phase has focused on concepts for annotating samples, environments, exposures, and measurements, the next phase will center on supporting annotation of results and findings, such as concept models of molecular, cellular and tissue effects. The value of the RBO will be determined in part by our ability to engage the community in its development, and we have established a Radiobiology Informatics Consortium with unrestricted membership as the owner of the RBO in order to encourage investigators, system owners and other to join in this effort. Anyone can report issues or request new concept modeling or other features directly on GitHub. By using the BioPortal application programming interface, systems can pose dynamic queries to the latest version of the RBO for information on individual classes or entire hierarchies; this design eliminates the need for systems to be updated in order to use newer versions of the RBO. We hope to contribute to the advancement of open radiobiological science through the continued, open development of the RBO, that will provide more precise, machine-interpretable descriptions of investigations, as well as support data meta-analysis through machine learning or other artificial intelligence methods.

knowledge↗

ProvLight: Efficient Workflow Provenance Capture on the Edge-to-Cloud Continuum

Modern scientific workflows require hybrid infrastructures combining numerous decentralized resources on the IoT/Edge interconnected to Cloud/HPC systems (aka the Computing Continuum) to enable their optimized execution. Understanding and optimizing the performance of such complex Edge-to-Cloud workflows is challenging. Capturing the provenance of key performance indicators, with their related data and processes, may assist in understanding and optimizing workflow executions. However, the capture overhead can be prohibitive, particularly in resource-constrained devices, such as the ones on the IoT/Edge.To address this challenge, based on a performance analysis of existing systems, we propose ProvLight, a tool to enable efficient provenance capture on the IoT/Edge. We leverage simplified data models, data compression and grouping, and lightweight transmission protocols to reduce overheads. We further integrate ProvLight into the E2Clab framework to enable workflow provenance capture across the Edge-to-Cloud Continuum. This integration makes E2Clab a promising platform for the performance optimization of applications through reproducible experiments.We validate ProvLight at a large scale with synthetic workloads on 64 real-life IoT/Edge devices in the FIT IoT LAB testbed. Evaluations show that ProvLight outperforms state-of-the-art systems like ProvLake and DfAnalyzer in resource-constrained devices. ProvLight is 26—37x faster to capture and transmit provenance data; uses 5—7x less CPU; 2x less memory; transmits 2x less data; and consumes 2—2.5x less energy. ProvLight [1] and E2Clab [2] are available as open-source tools.

Rosendo, Daniel↗