Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “containerization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Category V Compliant Container for Mars Sample Return Missions

A novel containerization technique that satisfies Planetary Protection (PP) Category V requirements has been developed and demonstrated on the mock-up of the Mars Sample Return Container. The proposed approach uses explosive welding with a sacrificial layer and cut-through-the-seam techniques. The technology produces a container that is free from Martian contaminants on an atomic level. The containerization technique can be used on any celestial body that may support life. A major advantage of the proposed technology is the possibility of very fast (less than an hour) verification of both containment and cleanliness with typical metallurgical laboratory equipment. No separate biological verification is required. In addition to Category V requirements, the proposed container presents a surface that is clean from any, even nonviable organisms, and any molecular fragments of biological origin that are unique to Mars or any other celestial body other than Earth.

Containers↗

Smart connected worker edge platform for smart manufacturing: Part 1—Architecture and platform design

Abstract The challenge of sustainably producing goods and services for healthy living on a healthy planet requires simultaneous consideration of economic, societal, and environmental dimensions in manufacturing. Enabling technology for data driven manufacturing paradigms like Smart Manufacturing (a.k.a. Industry 4.0) serve as the technological backbone from which sustainable approaches to manufacturing can be implemented. Unfortunately, these technologies are typically associated with broader and deeper factory automation that is often too expensive and complex for the small and medium sized manufacturers (SMMs) that comprise the majority of manufacturing business in the USA and for whom their most valuable asset are the people whose jobs automation while replace. This paper describes an edge intelligent platform to integrate internet‐of‐things technologies with computing hardware, software, computational workflows for machine learning, and data ingestion, enabling SMMs to transition into smart manufacturing paradigms by leveraging the intelligence of their people. The platform leverages consumer grade electronics and sensors (affordable and portable), customized software with open source software packages (accessible), and existing communication network infrastructures (scalable). The software systems are implemented via Kubernetes orchestration of Docker containerization to ensure scalability and programmability. The platform is adaptive via computational workflow engines that produce information from data by processing with low‐cost edge computing devices while efficiently accessing resources of cloud servers as needed. The proposed edge platform connects workers to technological resources that provide computational intelligence (i.e., silicon‐based sensing and computation for data collection and contextualization) to enable decision making at the edge of advanced manufacturing.

Kim, Yoon G.↗

Scalable Generation of High-fidelity Synthetic Population Ensembles

Used within social simulations, synthetic population ensembles enable uncertainty quantification (UQ) methods for obtaining more robust model inference and prediction. A synthetic population ensemble is a series of plausible virtual reconstructions of an area’s population at the granularity of people and residences, generated stochastically to preserve privacy of the source population survey’s respondents. In this paper, we demonstrate the production of large synthetic population ensembles for the U.S. via Oak Ridge National Laboratory’s UrbanPop framework to support modeling of high spatial resolution energy affordability metrics from nationwide social surveys in collaboration with the fusionACS project. The study involves two scenarios: creating ensembles for (1) 17 U.S. metropolitan areas in 2019 and (2) full U.S. Census Divisions in 2023, with each scenario consisting of 41 population instances (a base realization and 40 replicates). To accomplish this task at scale, we configured an integrated system within a research cloud, comprised of virtual containerizations, GPU-enhanced functionality, and orchestrated deployments of UrbanPop’s maturing Likeness Python ecosystem. Results demonstrate we maintained high-fidelity approximations of residential totals by areas of interest and the demographic characteristics of neighborhoods while reducing manual workflow burdens. Finally, we discuss plans to fine-tune and further develop our automated workflows for truly distributed job orchestration to increase computational efficiency, as well as provide an outlook for broadening applications of the ensembles.

Cluster computing↗

NWChem: Recent and Ongoing Developments

In this paper we summarize developments in the NWChem computational chemistry suite since the last major release (NWChem 7.0). Specifically, we focus on functionalities, along with input blocks, that are currently accessible in the current stable release (NWChem 7.2) and master branches, interfaces to quantum computing simulators, interfaces to external libraries, the NWChem GitHub repository, and containerization of NWChem executable images. In conclusion, some of the ongoing developments that will be available in the near future are also discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Packaging HEP Heterogeneous Mini-apps for Portable Benchmarking and Facility Evaluation on Modern HPCs

High Energy Physics (HEP) experiments are making increasing use of GPUs and GPU dominated High Performance Computer facilities. Both the software and hardware of these systems are rapidly evolving, creating challenges for experiments to make informed decisions as to where they wish to devote resources. In its first phase, the High Energy Physics Center for Computational Excellence (HEP-CCE) produced portable versions of a number of heterogeneous HEP mini-apps, such as p2r, FastCaloSim, Patatrack and the WireCell Toolkit, that exercise a broad range of GPU characteristics, enabling cross platform and facility benchmarking and evaluation. However, these miniapps still require a significant amount of manual intervention to deploy on a new facility. We present our work in developing turn-key deployments of these mini-apps, where by means of containerization and automated configuration and build techniques such as Spack, we are able to quickly test new hardware, software, environments and entire facilities with minimal user intervention, and then track performance metrics over time.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Exploiting Kubernetes to Simplify the Deployment and Management of the Multi-purpose CMS Pilot Job Factory

GlideinWMS, a widely utilized workload management system in high-energy physics (HEP) research, serves as the backbone for efficient job provisioning across distributed computing resources. It is utilized by various experiments and organizations, including CMS, OSG, Dune, and FIFE, to create HTCondor pools as large as 600k cores. In particular, a shared factory service historically deployed at UCSD has been configured to interface with more than 500 routes to compute clusters. As part of our team’s initiative to modernize infrastructure and enhance scalability, we undertook the migration of the GlideinWMS factory service into the Kubernetes environment. Leveraging the flexibility and orchestration capabilities of Kubernetes, we successfully deployed the factory service within the OSG Tiger Kubernetes cluster. The major benefits Kubernetes gives us is it streamlines the management and monitoring of the factory infrastructure, and improves fault tolerance through its resilient deployment strategies. Through this case study, we aim to share insights, challenges, and best practices encountered during the migration process. Our experience underscores the benefits of embracing containerization and Kubernetes orchestration for HEP computing infrastructure, paving the way for scalability and resilience in distributed computing environments.

Dost, Jeffrey Michael [UC, San Diego (main)]↗

GENTANGLE: integrated computational design of gene entanglements

The design of two overlapping genes in a microbial genome is an emerging technique for adding more reliable control mechanisms in engineered organisms for increased stability. The design of functional overlapping gene pairs is a challenging procedure, and computational design tools are used to improve the efficiency to deploy successful designs in genetically engineered systems. GENTANGLE (Gene Tuples ArraNGed in overLapping Elements) is a high-performance containerized pipeline for the computational design of two overlapping genes translated in different reading frames of the genome. This new software package can be used to design and test gene entanglements for microbial engineering projects using arbitrary sets of user-specified gene pairs.

59 BASIC BIOLOGICAL SCIENCES↗

SAIGE-GPU: accelerating genome- and phenome-wide association studies using GPUs

Genome-wide association studies (GWAS) at biobank scale are computationally intensive, especially for admixed populations requiring robust statistical models. SAIGE is a widely used method for generalized linear mixed-model GWAS but is limited by its CPU-based implementation, making phenome-wide association studies impractical for many research groups. We developed SAIGE-GPU, a GPU-accelerated version of SAIGE that replaces CPU-intensive matrix operations with GPU-optimized kernels. The core innovation is distributing genetic relationship matrix calculations across GPUs and communication layers. Applied to 2068 phenotypes from 635 969 participants in the Million Veteran Program, including diverse and admixed populations, SAIGE-GPU achieved a 5-fold speedup in mixed model fitting on supercomputing infrastructure and cloud platforms. We further optimized the variant association testing step through multi-core and multi-trait parallelization. Deployed on Google Cloud Platform and Azure, the method provided substantial cost and time savings. Source code and binaries are available for download at https://github.com/saigegit/SAIGE/tree/SAIGE-GPU-1.3.3. A code snapshot is archived at Zenodo for reproducibility (DOI: [10.5281/zenodo.17642591]). SAIGE-GPU is available in a containerized format for use across HPC and cloud environments and is implemented in R/C++ and runs on Linux systems.

Rodriguez, Alex [Argonne National Laboratory (ANL)↗

polishCLR: A Nextflow Workflow for Polishing PacBio CLR Genome Assemblies

Long-read sequencing has revolutionized genome assembly, yielding highly contiguous, chromosome-level contigs. However, assemblies from some third generation long read technologies, such as Pacific Biosciences (PacBio) continuous long reads (CLR), have a high error rate. Such errors can be corrected with short reads through a process called polishing. Although best practices for polishing non-model de novo genome assemblies were recently described by the Vertebrate Genome Project (VGP) Assembly community, there is a need for a publicly available, reproducible workflow that can be easily implemented and run on a conventional high performance computing environment. Here, we describe polishCLR (https://github.com/isugifNF/polishCLR), a reproducible Nextflow workflow that implements best practices for polishing assemblies made from CLR data. PolishCLR can be initiated from several input options that extend best practices to suboptimal cases. It also provides re-entry points throughout several key processes, including identifying duplicate haplotypes in purge_dups, allowing a break for scaffolding if data are available, and throughout multiple rounds of polishing and evaluation with Arrow and FreeBayes. PolishCLR is containerized and publicly available for the greater assembly community as a tool to complete assemblies from existing, error-prone long-read data.

59 BASIC BIOLOGICAL SCIENCES↗

Unveiling the microbial realm with VEBA 2.0: a modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic and viral multi-omics from either short- or long-read sequencing

Abstract The microbiome is a complex community of microorganisms, encompassing prokaryotic (bacterial and archaeal), eukaryotic, and viral entities. This microbial ensemble plays a pivotal role in influencing the health and productivity of diverse ecosystems while shaping the web of life. However, many software suites developed to study microbiomes analyze only the prokaryotic community and provide limited to no support for viruses and microeukaryotes. Previously, we introduced the Viral Eukaryotic Bacterial Archaeal (VEBA) open-source software suite to address this critical gap in microbiome research by extending genome-resolved analysis beyond prokaryotes to encompass the understudied realms of eukaryotes and viruses. Here we present VEBA 2.0 with key updates including a comprehensive clustered microeukaryotic protein database, rapid genome/protein-level clustering, bioprospecting, non-coding/organelle gene modeling, genome-resolved taxonomic/pathway profiling, long-read support, and containerization. We demonstrate VEBA’s versatile application through the analysis of diverse case studies including marine water, Siberian permafrost, and white-tailed deer lung tissues with the latter showcasing how to identify integrated viruses. VEBA represents a crucial advancement in microbiome research, offering a powerful and accessible software suite that bridges the gap between genomics and biotechnological solutions.

59 BASIC BIOLOGICAL SCIENCES↗

Network Anomaly Detection in Distributed Edge Computing Infrastructure

As networks continue to grow in complexity and scale, detecting anomalies has become increasingly challenging, particularly in diverse and geographically dispersed environments. Traditional approaches often struggle with managing the computational burden associated with analyzing large-scale network traffic to identify anomalies. This paper introduces a distributed edge computing framework that integrates federated learning with Apache Spark and Kubernetes to address these challenges. We hypothesize that our approach, which enables collaborative model training across distributed nodes, significantly enhances the detection accuracy of network anomalies across different network types. We show that by leveraging distributed computing and containerization technologies, our framework not only improves scalability and fault tolerance but also achieves superior detection performance compared to state-of-the-art methods. Extensive experiments on the UNSW-NB15 and ROAD datasets validate the effectiveness of our approach, demonstrating statistically significant improvements in detection accuracy and training efficiency over baseline models, as confirmed by MannWhitney U and Kolmogorov-Smirnov tests (p<0.05).

Marfo, William [University of Texas at El Paso,Dep↗

VAC: A Software Approach to Resilient SCADA Automation

To better secure critical infrastructure, especially power systems, this paper introduces a virtual SCADA automation controller. The automation controller is a gateway into a power subsystem, making it a valuable target for cyber-attacks that could cut it off from the control center and cause a loss of view and control. To prevent this, the Virtual Automation Controller (VAC) is a backup device that mirrors the capabilities of the physical controller. It can communicate via Modbus and DNP3 and is containerized so it can be deployed on a variety of platforms. Furthermore, it utilizes software-defined networking to quickly disconnect a failed automation controller and preserve its state for forensics. The VAC gives system operators time to replace the failed controller and prevents dangerous and costly damage to power systems. The VAC is compared against the SEL 3505-3 RTAC and shown to have the necessary features to act as a failover controller.

Johnson, Jordan↗

Virtual Framework for Development and Testing of Federation Software Stack

Softwarization of networked infrastructures combined with containerization of codes promises unprecedented computing capabilities distributed across the federations of computing systems and physical instruments. The development and testing of a software stack that implements these capabilities over an expensive physical production infrastructure is not cost-effective, and in the early stages, may potentially cause service disruptions. To address these aspects, we develop the Virtual Federated Science Instrument Environment (VFSIE), a digital twin of the physical infrastructure that emulates a multi-site federation. Each federated site is emulated using containers and virtual hosts that are connected over local-area networks, and the sites, in turn, are connected over an emulated wide-area network. We describe the framework design and implementation details. We also illustrate its application by emulating a federation of four laboratories that use Jupyter Notebook for computations and the EPICS software system for instrument control.

Al Najjar, Anees↗

Deciphering Discrepancies: A Comparative Analysis of Docker Image Security

As the use of microservices continues to grow and become a foundational approach to architecting software solutions, ensuring the security of microservices is paramount. Docker images have emerged as the predominant solution to containerize microservices–and thus, Docker images are becoming a large attack surface. Thus, reducing vulnerabilities in Docker images will reduce microservice cyberattacks. A common way to find vulnerabilities in Docker images employs static analysis tools like Trivy and Grype. However, these tools frequently generate disparate vulnerability reports when analyzing the same Docker image, thus causing uncertainty in tool selection. We collected 927 Docker images, analyzed them with Trivy and Grype, and compared the vulnerabilities reported in each image. Among the 865 images found to have vulnerabilities, Trivy and Grype disagreed on both the number of vulnerabilities and the vulnerability IDs found therein. Since both tools interface with external vulnerability databases, some discrepancies can be attributed to how the tools interface with these external resources. The external vulnerability databases partially overlap and frequently contradict one another, thereby creating challenges for static analysis tool developers and end users alike. This New Ideas and Emerging Results (NIER) study contains new and critical information that practitioners need for selecting and using static analysis tools–given that increases in the use of Docker technologies means increases in the size of the attack surfaces.

Boles, Brittany [Montana State University]↗

Gene Tuples ArraNGed in overLapping Elements

GENTANGLE is a high performance containerized pipeline for the computational design of two overlapping genes translated in different reading frames of the genome that can be used to design and test gene entanglements for new microbial engineering projects using arbitrary sets of user specified gene pairs.

Leonard, SeanP↗

Productivity frameworks for HPC

Productivity Frameworks for HPC will include container recipes, build recipes, continuous integration scripts, and other software aimed at testing the portability of containerized HPC software across platforms and interconnects. In particular, it tests the utility of bind-mounting at the MPI layer (rather than the underlying fabric layer) to leverage a standardized protocol and avoid various technical debt and vendor lock-in. Since MPIs are often ABI-incompatible, trampolines such as the open-source Wi4MPI will be tested when such cases arise.

Hanford, Nathan↗

Platform for Integrated Land use And Transportation Experiments and Simulation (PILATES) v1.0

PILATES allows for flexibly and at-scale coupling of multiple models to allow for multi-scale and multi-resolution simulation of regional-scale transport networks. In particular, it couples the MATSim-derived transportation modeling framework for Behavior, Energy, Autonomy and Mobility (BEAM) with other models operating at different time scales. Rather than tightly coupling supply and demand models using shared agents and memory within the same software process, PILATES orchestrates different model runs in a containerized framework. This structure requires passing information from the demand models to BEAM in the format of a synthetic population and agent plans, and from BEAM to the demand models in terms or origin/destination tables (also known as "skims"). This allows it to take advantage of the behavioral sophistication of existing activity-based models as well as the reinforcement learning structure of MATSim replanning and adopted by BEAM, in a way that requires minimal changes to existing models. It also takes advantage of the computational performance of BEAM, which allows for simulations with millions of agents to complete in reasonable time as well as allowing for detailed mechanistic simulation of the operation of on-demand modes.

Needell, Zachary↗

Janus v1.0

Janus provides a software framework for lightweight container management and orchestration. It's primary use cases are around deploying containerized services for high-performance data movement needs. Thus, Janus differentiates itself from systems like Kubernetes by tailoring the deployment of containers around network, storage, and host tuning optimizations. Janus uses the concept of profiles to capture repeatable deployment patterns and applies them to container execution across one or more endpoints. A Janus Agent component provides remote host tuning and monitoring capabilities.

Essiari, Abdelilah [Lawrence Berkeley National Lab↗