Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Jobs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

Elastic Resource Management for Deep Learning Applications in a Container Cluster

The increasing demand for learning from massive datasets is restructuring our economy. Effective learning, however, involves nontrivial computing resources. Most businesses utilize commercial infrastructure providers (e.g., AWS) to host their computing clusters in the cloud, where various jobs compete for available resources. While cloud resource management is a fruitful research field that has made many advances in production, such as Kubernetes and YARN, few efforts have been invested to further optimize the system performance, especially for deep learning (DL) training jobs in a container cluster. This work introduces FlowCon, a system that is able to monitor the individual evaluation functions of DL jobs at runtime, and thus to make placement decisions on resource allocations elastically. Here, we present a detailed design and implementation of FlowCon and conduct intensive experiments over various DL models. The results demonstrate that FlowCon significantly improves DL job completion time and resource utilization efficiency, compared to default systems. According to the results, FlowCon is able to improve the completion time by up to 68.8% and meanwhile, reduce the makespan by 18.0%, in the presence of various DL job workloads.

97 MATHEMATICS AND COMPUTING↗

Graph neural networks for detecting anomalies in scientific workflows

Identifying and addressing anomalies in complex, distributed systems can be challenging for reliable execution of scientific workflows. We model these workflows as directed acyclic graphs (DAGs), where the nodes and edges of the DAGs represent jobs and their dependencies, respectively. We develop graph neural networks (GNNs) to learn patterns in the DAGs and to detect anomalies at the node (job) and graph (workflow) levels. We investigate workflow-specific GNN models that are trained on a particular workflow and workflow-agnostic GNN models that are trained across the workflows. Our GNN models, which incorporate both individual job features and topological information from the workflow, show improved accuracy and efficiency compared to conventional learning methods for detecting anomalies. While joint trained with multiple scientific workflows, our GNN models reached an accuracy more than 80% for workflow level and 75% for job level anomalies. In addition, we illustrate the importance of hyperparameter tuning method in our study that can significantly improve the metric(s) measure of evaluating the GNN models. Finally, we integrate explainable GNN methods to provide insights on job features in the workflow that cause an anomaly.

97 MATHEMATICS AND COMPUTING↗

BEE - FY20 P6-3: Release BEEWorkflowManager, BEETaskManager, and client application 2.3.6.01 – LANL ATDM ST / STNS01-4 P6 Milestone Completion Documentation

Release BEEWorkflowManager, BEETaskManager, and client software. The BEEWorkflowManager daemon runs on the HPC cluster login node. It accepts workflows submitted by the BEE client. These workflows are specified using the Common Workflow Language (CWL) standard. The BEEWorkflowManager loads workflows into the Neo4j graph database to create the workflow directed acyclic graph (DAG), and submits the workflow tasks to the BEETaskManager for execution. The BEEWorkflowManager records the state of the workflow and its tasks, and communicates this state to the BEE client. The BEEWorkflowManager will start, pause, and cancel a running workflow and its tasks at the command of the BEE client. The BEETaskManager daemon runs on the HPC cluster login node. It accepts tasks from the BEEWorkflowManager, turns those tasks into HPC resource manager jobs (e.g. a slurm job script), and submits the job to the cluster resource manager. The BEETaskManager then tracks the status of the job (pending, running, complete) and updates the BEEWorkflowManager. The BEETaskManager will also cancel a queued or running job when commanded to do so by the BEEWorkflowManager. The first release of the BEETaskManager will support the Slurm resource manager and the Charliecloud linux container runtime.

97 MATHEMATICS AND COMPUTING↗

Creating Unit Tests for GlideinWMS using AI tools

GlideinWMS is a workload management system that uses distributed computing to complete tasks, also known as jobs. It is particularly useful for high-throughput computing that’s used in research projects. It relies on Glideins, which are pilot jobs that pull jobs from a queue and provide resources for their completion, based on the jobs requirements. These decisions are made based on resource availability and job requirements. We used new AI tools to add unit tests to GlideinWMS.

Baburashvili, Ilya↗

Fermilab s Transition to Token Authentication

Fermilab is the first High Energy Physics institution to transition from X.509 user certificates to authentication tokens in production systems. All of the experiments that Fermilab hosts are now using JSON Web Token (JWT) access tokens in their grid jobs. Many software components have been either updated or created for this transition, and most of the software is available to others as open source. The tokens are defined using the WLCG Common JWT Profile. Token attributes for all the tokens are stored in the Fermilab FERRY system which generates the configuration for the CILogon token issuer. High security-value refresh tokens are stored in Hashicorp Vault configured by htvault-config, and JWT access tokens are requested by the htgettoken client through its integration with HTCondor. The Fermilab job submission system jobsub was redesigned to be a lightweight wrapper around HTCondor. For automated job submissions a managed tokens service was created to reduce duplication of effort and knowledge of how to securely keep tokens active. The existing Fermilab file transfer tool ifdh was updated to work seamlessly with tokens, as well as the Fermilab POMS (Production Operations Management System) which is used to manage automatic job submission and the RCDS (Rapid Code Distribution System) which is used to distribute analysis code via the CernVM FileSystem. The dCache storage system was reconfigured to accept tokens for authentication in place of X.509 proxy certificates. As some services and sites have not yet implemented token support, proxy certificates are still sent with jobs for backwards compatibility but some experiments are beginning to transition to stop using them. There have been some glitches and learning curve issues but in general the system has been performing well and is being improved as operational problems are addressed.

Dykstra, David↗

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

Accelerating Machine Learning Inference with GPUs in ProtoDUNE Data Processing

Abstract We study the performance of a cloud-based GPU-accelerated inference server to speed up event reconstruction in neutrino data batch jobs. Using detector data from the ProtoDUNE experiment and employing the standard DUNE grid job submission tools, we attempt to reprocess the data by running several thousand concurrent grid jobs, a rate we expect to be typical of current and future neutrino physics experiments. We process most of the dataset with the GPU version of our processing algorithm and the remainder with the CPU version for timing comparisons. We find that a 100-GPU cloud-based server is able to easily meet the processing demand, and that using the GPU version of the event processing algorithm is two times faster than processing these data with the CPU version when comparing to the newest CPUs in our sample. The amount of data transferred to the inference server during the GPU runs can overwhelm even the highest-bandwidth network switches, however, unless care is taken to observe network facility limits or otherwise distribute the jobs to multiple sites. We discuss the lessons learned from this processing campaign and several avenues for future improvements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Employment access assessed using the mobility energy productivity (MEP) metric

Transit agencies, local governments, employers, and job-seekers have a shared interest in connecting residents with jobs in an affordable and time efficient manner, with public agencies also caring about energy efficiency and air quality. Employment hubs are an opportunity to solve the spatial mismatch between homes of job-seekers and the locations of desirable jobs. Such is the case in Columbus, Ohio, between Rickenbacker Industrial Park and the Linden neighborhood, which has experienced persistent poverty. Using the mobility energy productivity (MEP) metric to examine travel time, cost, and energy efficiency, we show current transit service is undesirable due to excessive travel time (70 min, MEP = 0), while driving alone (MEP = 0.20) may be less desirable than a hypothetical, fare-free microtransit service (MEP = 0.23). Updating MEP to use locally-derived input data can help identify parameters under which providing microtransit service in a specific place has compelling benefits in terms of vehicle energy efficiency as well as cost and travel time for riders.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Intersections of Disadvantaged Communities and Renewable Energy Potential: Data Set and Analysis to Inform Equitable Investment Prioritization in the United States

Renewable energy development can bolster local economies through job creation, local tax revenues, and reduced energy costs; however, communities most in need of economic development and employment opportunities often see lower levels of renewable energy deployment. We sought to identify areas where disadvantaged community indicators and high generation potential from cost-effective renewable energy opportunities intersect and deployment could lead to economic development and job creation. This presentation will highlight several of our findings. This research and the associated county-level data set are intended to inform national- and state-level energy-related assistance programs, economic development efforts, and infrastructure programs seeking to prioritize investments in disadvantaged communities.

community energy planning↗

Automated and Distributed Monte Carlo Generation for GlueX

MCwrapper is a set of systems that manages the entire Monte Carlo production workflow for GlueX and provides standards for how that Monte Carlo is produced. MCwrapper was designed to be able to utilize a variety of batch systems in a way that is relatively transparent to the user, thus enabling users to quickly and easily produce valid simulated data at home institutions worldwide. Additionally, MCwrapper supports an autonomous system that takes user’s project submissions via a custom web application. The system then atomizes the project into individual jobs, matches these jobs to resources, and monitors the jobs status. The entire system is managed by a database which tracks almost all facets of the systems from user submissions to the individual jobs themselves. Users can interact with their submitted projects online via a dashboard or, in the case of testing failure, can modify their project requests from a link contained in an automated email. Beginning in 2018 the GlueX Collaboration began to utilize the Open Science Grid (OSG) to handle a bulk of simulation tasks; these tasks are currently being performed on the OSG automatically via MCwrapper. This talk will outline the entire system of MCwrapper, its use cases, and the unique challenges facing the system.

Britton, Thomas↗

A Digital Twin of Scalable Quantum Clouds

Quantum computing has emerged as a transformative technology capable of solving complex problems beyond the limit of classical systems. The rapid development of quantum processors has led to the proliferation of cloud-based quantum computing services offered by platforms such as IBM, Google, and Amazon. These platforms introduce unique challenges in resource allocation, job scheduling, and multi-device orchestration as quantum workloads become increasingly complex. In this work, we present a digital twin of quantum cloud infrastructures: a framework designed to model and simulate the behavior of real quantum cloud systems. Developed in Python using the SimPy discrete-event simulation library, the framework replicates key aspects of quantum cloud environments, including detailed quantum device modeling, job lifecycle management, and job fidelity. It incorporates noise-aware fidelity estimation, making it the first of its kind to simulate superconducting gate-based quantum cloud systems at an administrative level with job fidelity. We present use cases as proof of concept, demonstrating that our quantum cloud simulation framework can act as a digital twin of a quantum cloud and support the modeling and implementation of practical systems.

Luo, Waylon [Kent State University]↗

Profiling the BLAST bioinformatics application for load balancing on high-performance computing clusters

Abstract Background The Basic Local Alignment Search Tool (BLAST) is a suite of commonly used algorithms for identifying matches between biological sequences. The user supplies a database file and query file of sequences for BLAST to find identical sequences between the two. The typical millions of database and query sequences make BLAST computationally challenging but also well suited for parallelization on high-performance computing clusters. The efficacy of parallelization depends on the data partitioning, where the optimal data partitioning relies on an accurate performance model. In previous studies, a BLAST job was sped up by 27 times by partitioning the database and query among thousands of processor nodes. However, the optimality of the partitioning method was not studied. Unlike BLAST performance models proposed in the literature that usually have problem size and hardware configuration as the only variables, the execution time of a BLAST job is a function of database size, query size, and hardware capability. In this work, the nucleotide BLAST application BLASTN was profiled using three methods: shell-level profiling with the Unix “time” command, code-level profiling with the built-in “profiler” module, and system-level profiling with the Unix “gprof” program. The runtimes were measured for six node types, using six different database files and 15 query files, on a heterogeneous HPC cluster with 500+ nodes. The empirical measurement data were fitted with quadratic functions to develop performance models that were used to guide the data parallelization for BLASTN jobs. Results Profiling results showed that BLASTN contains more than 34,500 different functions, but a single function, RunMTBySplitDB, takes 99.12% of the total runtime. Among its 53 child functions, five core functions were identified to make up 92.12% of the overall BLASTN runtime. Based on the performance models, static load balancing algorithms can be applied to the BLASTN input data to minimize the runtime of the longest job on an HPC cluster. Four test cases being run on homogeneous and heterogeneous clusters were tested. Experiment results showed that the runtime can be reduced by 81% on a homogeneous cluster and by 20% on a heterogeneous cluster by re-distributing the workload. Discussion Optimal data partitioning can improve BLASTN’s overall runtime 5.4-fold in comparison with dividing the database and query into the same number of fragments. The proposed methodology can be used in the other applications in the BLAST+ suite or any other application as long as source code is available.

59 BASIC BIOLOGICAL SCIENCES↗

April 2020 Darshan counters from the Summit supercomputer

This dataset is the Darshan counters collected from the Summit supercomputer in a month of April 2020. 1. Description of methods used for collection/generation of data: Job submitted on Summit HPC system when completed successfully and has made I/O calls (captured by Darshan tool) writes a Darshan log file on alpine filesystem. One job can have multiple `jsrun` commands and Darshan will generate separate logs each log corresponding to an `jsrun` command, so a job can have one or more Darshan logs associated with it. 2. Methods for processing the data: To process the data, we first use `darshan-util` tool to parse the Darshan logs. Then we restructure the logs and merge data from multiple Darshan logs if they belong to the same Summit job.

97 MATHEMATICS AND COMPUTING↗

Tribal Renewable Energy: Bishop Paiute Tribe, Residential Solar Program Phase II (Final Report)

The project consisted of the design, installation, inspection, and interconnection of 35 grid-tied, solar electric systems, totaled 123 kW rated capacity, on qualified existing low-income single-family homes located within the Bishop Paiute Reservation, which provide at least 65-79% savings in displaced electricity. Tribal job trainees were hired for each installation, gaining valuable job experience. Additionally, each homeowner was educated on energy efficiency and renewable energy. The overall project goal was to deploy clean energy systems in order to achieve the Bishop Paiute Tribe’s long-term goals of energy self-sufficiency, environmental protection, and better lives for our Tribal members and community. The energy displaced was 82.4 percent of total electricity used (exceeding estimated 65-79%)or over 170,000 kWh/year generating approximately $1.29 million worth of power for low-income families over their lifespans, while eliminating an estimated 2,640 tons of greenhouse gas emissions. This reduction makes a significant difference on the reservation giving more families money to spend on other essential items, while reducing their carbon footprint in this beautiful mountain community. Also, implementing these 35 systems provided a approximately 132 hours of paid solar installation work for tribal members. Training and good paying jobs are scarce on the reservation and solar is the fastest growing industry in CA and this training offered our members a real chance to learn and then get paid. The program is very significant to the family who qualified, as most families fall below the federal poverty guidelines, and many are living paycheck to paycheck. Overall, the triple impact of the Bishop Paiute Tribe Residential Solar Program Phase II— affordable energy for low-income families, on-site clean energy production, and hands-on solar installation jobs for local workers—these all help build the Tribe’s energy, economic, environmental, and social self-sufficiency and sovereignty amongst the most needy on the Reservation

14 SOLAR ENERGY↗

Cost-Benefit Analysis For Indonesia Building Sector: Whole-Building Cooling Solutions

The Net Zero World (NZW) Initiative Collaborative Work Program with the Government of Indonesia (GoI) includes technical assistance and investment mobilization facilitation to accelerate deployment of energy efficiency technologies and solutions for the building sector. A February 2023 U.S.–Indonesia Joint Workshop on Decarbonizing the Building Sector yielded a NZW Indonesia Building Decarbonization Working Group (NZW IBDWG) with four sub-working groups (SWG): SWG-A National Center, SWG-B Capacity Building, SWG-C Investment and Financing, and SWG-D Pilot Projects. Technical analysis of whole-building cooling solutions for tropical climates of Indonesia was conducted by SWG-A to quantify energy savings, carbon dioxide reductions, and comfort improvements offered by 12 passive or low-energy cooling strategies: ceiling fans with and without thermostat setbacks; cool roofs; cool walls; exterior awnings; exterior shades; interior shades; insulated roofs; insulated walls; low-e windows; solar window films; and natural ventilation. Leveraging the results from SWG-A, cost-benefit analysis (CBA) was conducted by SWG-C to assess the consumer and national costs and impacts associated with these 12 cooling solutions. The evaluation involved estimating life-cycle costs (LCC), payback period (PBP), net present values (NPV), annual electricity burden change for low-income households, and reduced national annual power-sector generation demand by 2030, 2040, 2050, and 2060. This evaluation can help guide Indonesia’s Just Energy Transition Partnership (JETP) investments in policies and programs to advance research, development, deployment, and commercial adoption (RDDCA) of efficient residential building sector cooling technologies and solutions in Indonesia. Four key energy conservation measures (ECM) have been identified to reduce air-conditioning (AC) energy demand in single-family housing in Indonesia: ceiling fan with temperature setback (to 28.1 °Celcius from 25 °C); insulated walls; insulated roof; and cool roof. This study found that low-income households with AC installations in Indonesia currently face a high energy cost burden of approximately 10%. However, by implementing a ceiling fan with temperature setback, this burden could decrease to 2.5% today and further reduce to 1.3% by the year 2060. The PBP for a ceiling fan with temperature setback is one year, indicating one of the lowest LCC and best NPV. In the planned upcoming phase of CBA, a series of building cooling improvement scenarios can be further defined, incorporating more than one ECM in combination with socio-economic factors evaluated in the initial CBA phase. Additionally, the analysis of ECM effects in multifamily housing can be expanded. This broader national analysis aims to encompass a holistic and comprehensive system-level perspective, including factors such as avoided power sector infrastructure investments, domestic job creation, domestic manufacturing job creation, and gross domestic product (GDP) growth.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A Mid-Century Net-Zero Scenario for the State of Wyoming and its Economic Impacts

Clean hydrogen has the potential to help achieve 10% economy-wide emissions reductions by 2050 relative to 2005, promote energy security and resilience, and develop a new economy in the United States. In 2030, the hydrogen economy could create about 100,000 new jobs to build new capital projects and clean hydrogen infrastructure. The Wyoming Energy Authority recently announced the state’s energy strategy, which establishes a goal of net-zero emissions by 2050. Under all likely scenarios, achieving a mid-century net-zero target will pose challenges and create opportunities for Wyoming’s energy sector. If executed properly, the transition could favorably affect the state’s economy overall in the long term. This research program examines the economic impact of fossil energy production in Wyoming and provides various predictions for future energy mixes to achieve net-zero emissions. Preliminary work suggests that Wyoming-based hydrogen production could have significant economic benefits and job creation implications for Wyoming. This study further assesses Wyoming’s opportunities to create hydrogen-based industries, assess economic impacts, identify knowledge gaps and research needs, and create a Hydrogen Center of Excellence to accelerate commercialization and deployment. This project helped to understand Wyoming's areas of focus for research and development and identified its areas of strength and potential challenges in creating a hydrogen ecosystem. As a result of this study, we estimate that for blue hydrogen produced from coal and gas resources, the overall cost reduction will be driven mainly by the carbon-sequestration tax credit and the improvement in carbon capture. Mature technologies, like SMR and PSA, will make limited contributions. They have no or limited reductions from an additional capacity deployment in future costs. We also understand the importance of continued support from public and private sectors for Carbon Capture and Storage (CCS)-related research, development, and demonstration programs at federal and state levels. The successful and efficient production of blue hydrogen requires a unique blend of energy resources, geology, regulation, law, and infrastructure. Wyoming has the distinction of meeting all these demands. The team also estimates that the availability and command of water resources accessible for hydrogen production are crucial for developing new projects. Water treatment, use, and disposal after treatment will also make projects possible. Primarily, this is relevant for hydrogen made using renewable energy. Wyoming has one of the best wind resource capacity in the nation. Harnessing this resource is challenging due to limited transmission line availability. Hydrogen could become one of the solutions to the stranded resource problem, primarily if the water availability challenge is addressed. Using produced oil & gas water could help to solve the problem. A commonly cited barrier to the expansion of hydrogen markets is the cost associated with constructing new pipelines, which typically require large amounts of capital to develop. Wyoming already possesses much of the export infrastructure needed to connect Wyoming’s hydrogen production with major markets across the West Coast, Pacific Northwest, Midwest, and Front Range regions of the United States, where a large portion of Wyoming’s natural gas is already transported. In addition to transportation by pipeline, rail transportation of hydrogen has also proven feasible. Wyoming uses its extensive railway system to transport large amounts of coal to its export partners across the United States. By using cryogenic or compressed-gas cars, Wyoming has the potential to add hydrogen to its existing network of railroad energy exports. The same technology may also be applied to hydrogen transport via trucks traveling interstate highways. Wyoming’s workforce is ready to meet the demands of clean hydrogen development. Many of the skills and training needed for hydrogen production are the same skills already possessed by Wyoming’s oil & gas and coal workforce. Many government and industry leaders expect clean hydrogen and other low-carbon energy projects to generate significant job growth and to recruit many already-trained oil & gas and coal workers whose jobs may be displaced. As energy companies seek to penetrate the markets for Wyoming hydrogen production, there is a natural mutual benefit to Wyoming’s workers and companies seeking to launch projects with the assistance of a trained workforce. Wyoming’s university and community college system have adopted several programs to ensure that highly qualified engineers and other technically skilled employees continue to graduate with skills to support the development of hydrogen and other innovative energy projects moving forward. Throughout the project, stakeholder outreach and education took many forms, including meetings with several major companies in the industry, collaborating with local government organizations, educational organizations, and national laboratories, tribal outreach and engagement, the sponsoring of several hydrogen-focused projects in many departments throughout the University of Wyoming, and developing a collaboration with international universities. The products of these collaborations consist of working relationships with several companies in the industry, educational institutions, national labs, and local government, as well as strong connections with individuals who will play an essential role in the success of the Hydrogen Energy Research Center.

08 HYDROGEN↗

A Managed Tokens Service for Securely Keeping and Distributing Grid Tokens

Fermilab is transitioning authentication and authorization for grid operations to using bearer tokens based on the WLCG Common JWT (JSON Web Token) Profile. One of the functionalities that Fermilab experimenters rely on is the ability to automate batch job submission, which in turn depends on the ability to securely refresh and distribute the necessary credentials to experiment job submit points. Thus, with the transition to using tokens for grid operations, we needed to create a service that would obtain, refresh, and distribute tokens for experimenters’ use. This service would avoid the need for experimenters to be experts in obtaining their own tokens and would better protect the most sensitive long-lived credentials. Further, the service needed to be widely scalable, as we are currently keeping credentials active for approximately 15 experiments, each with 1-3 different credentials, and distributing those credentials to 2-20 submit points per experiment, with those numbers steadily increasing. To address these issues, we created and deployed a Managed Tokens service. The service is written in Go, taking advantage of that language’s native concurrency primitives to easily be able to scale operations as we onboard experiments. The service uses as its first credentials a set of kerberos keytabs, stored on the same secure machine that the Managed Tokens service runs on. These kerberos credentials allow the service to use htgettoken via condor_vault_storer to store vault tokens in the HTCondor credential managers (credds) that run on the batch system scheduler machines (HTCondor schedds); as well as downloading a local, shorter-lived copy of the vault token. The kerberos credentials are then also used to distribute copies of the locally-stored vault tokens to experiment submit points. When experimenters schedule jobs to be submitted, these distributed vault tokens are used to access a Hashicorp Vault instance (run separately from the Managed Tokens service), and previously-stored refresh tokens there are used to obtain the bearer token that is submitted with the job. We will discuss here the design of the Managed Tokens service, including elaborating on certain choices we made with regards to concurrent operations, configuration, monitoring, and deployment.

Bhat, Shreyas↗