Virtual Infrastructure Twin for Computing-Instrument Ecosystems: Software and Measurements
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Explore the source record for details and available documents.
Collaboration and team science are emerging areas of interest in software production. Historically, multi-institutional research collaborations are difficult to initiate and maintain, negatively impacting communication, negotiation, and dialogue between industry, government, and academic researchers. The Exascale Computing Project (ECP), a massive, multi-team, high-stakes initiative, facilitated broader research collaboration under a shared funding structure and extended timeline to support scientific discovery. Here, we conducted interviews with ECP teams, representing a variety of domain specialties, research institutions, and programming backgrounds. Using thematic analysis, we assessed how ECP’s structure created an environment of increased trust among projects and how software shared between teams facilitated sustained collaboration. We found that the expectation of future collaboration, i.e., the shadow of the future, greatly enhanced trust among teams and the quality of scientific software produced. Based on our findings within ECP projects, we connect to the existing literature on trust in software engineering and share recommendations for sustainable multi-institutional collaboration and shared best software practices.
The CABA project is a test of a new tool, Survey,developed by Trenza to be used not only for benchmarking or profiling programs but also to allow incorporation of the information provided by Survey to be utilized in a CI, continuous integration,tool such as GitLab CI.Survey is foremost a means of assessing code performance in terms of time and operations which for computer programmers is known as benchmarking.
This document describes Offline Software and Computing for the Deep Underground Neutrino Experiment (DUNE) experiment, in particular, the conceptual design of the offline computing needed to accomplish its physics goals. Our emphasis in this document is the development of the computing infrastructure needed to acquire, catalog, reconstruct, simulate and analyze the data from the DUNE experiment and its prototypes. In this effort, we concentrate on developing the tools and systems that facilitate the development and deployment of advanced algorithms. Rather than prescribing particular algorithms, our goal is to provide resources that are flexible and accessible enough to support creative software solutions as HEP computing evolves and to provide computing that achieves the physics goals of the DUNE experiment.
DUNE, like other HEP experiments, faces a challenge related to matching execution patterns of our production simulation and data processing software to the limitations imposed by modern high-performance computing facilities. In order to efficiently exploit these new architectures, particularly those with high CPU core counts and GPU accelerators, our existing software execution models require adaptation. In addition, the large size of individual units of raw data from the far detector modules pose an additional challenge somewhat unique to DUNE. Here we describe some of these problems and how we begin to solve them today with existing software frameworks and toolkits. We also describe ways we may leverage these existing software architectures to attack remaining problems going forward. This whitepaper is a contribution to the Computational Frontier of Snowmass21.
Explore the source record for details and available documents.
In nuclear physics (NP) today the study of quarks, gluons and their strong interactions extends across a broad research program at a varied range of collaborative scales, from a few collaborators up to large experiments at scales comparable to those typical of high energy physics (HEP). Overall, the software and computing efforts vary accordingly, from pragmatic do-it-yourself approaches among a few, to substantial organized software and computing activities within large experiments. With new experiments starting up and on the horizon [1], and rapidly increasing data volumes [2, 3] and processing demands even at small experiments, the NP community has in recent years been thinking about the next generation of data processing and analysis workflows that will maximize the science output. One context for this discussion has been a series of workshops, “Future Trends in Nuclear Physics Computing” [4]. The most recent in this series took place in Fall 2020, organized by the authors together with colleagues. The workshop focused on identifying the unique aspects of software and computing in NP, and discussing how the NP community could strengthen common efforts and chart a path forward for the next decade, sure to be an exciting one with rich ongoing scientific programs at Brookhaven National Laboratory (BNL), Jefferson Lab (JLab), and other NP facilities, and culminating in datataking at the Electron-Ion Collider (EIC) [5,6,7] in the early 2030s. Without claiming to present a collective view from the workshop and discussions since—fortunately this is not expected of us in this opinion editorial—we offer here our reflections on the topic, informed by the workshop and the summary we authored with our colleagues [8], as well as discussions and developments in the eventful time since.
The Electron Ion Collider (EIC) is the next generation of precision QCD facility to be built at Brookhaven National Laboratory in conjunction with Thomas Jefferson National Laboratory. There are a significant number of software and computing challenges that need to be overcome at the EIC. During the EIC detector proposal development period, the ECCE consortium began identifying and addressing these challenges in the process of producing a complete detector proposal based upon detailed detector and physics simulations. Here, in this document, the software and computing efforts to produce this proposal are discussed; furthermore, the computing and software model and resources required for the future of ECCE are described.
Software that orchestrate data processing and management for mass spectrometry workflows and data products. Stack presented contains a web application, data processing job scheduling, and data processing workers for data processing and management.
The demand for high performance computing (HPC) resources continues to grow, driven by the increasing complexity of modeling and simulation, artificial intelligence (AI), and machine learning (ML) workloads [Porter]. The growing energy consumption demand of these HPC systems is a significant concern, both in terms of operational costs and environmental impact. AI hardware accelerators are expected to reach 1.5% of the world’s power consumption by 2029 [Shah].
We summarize the status of Deep Underground Neutrino Experiment (DUNE) Offline Software and Computing program. We describe plans for the computing infrastructure needed to acquire, catalog, reconstruct, simulate and analyze the data from the DUNE experiment and its prototypes in pursuit of the experiment's physics goals of precision measurements of neutrino oscillation parameters, detection of astrophysical neutrinos, measurement of neutrino interaction properties and searches for physics beyond the Standard Model. In contrast to traditional HEP computational problems, DUNE's Liquid Argon Time Projection Chamber data consist of simple but very large (many GB) data objects which share many characteristics with astrophysical images. We have successfully reconstructed and simulated data from 4% prototype detector runs at CERN. The data volume from the full DUNE detector, when it starts commissioning late in this decade will present memory management challenges in conventional processing but significant opportunities to use advances in machine learning and pattern recognition as a frontier user of High Performance Computing facilities capable of massively parallel processing. Our goal is to develop infrastructure resources that are flexible and accessible enough to support creative software solutions as HEP computing evolves.
To manage the complex demands of modern high-performance computing (HPC), software applications increasingly depend on software developed by other teams, often at other institutions. An HPC software ecosystem approach is required to support dependencies on third-party scientific software. An ecosystem approach provides layers of activity above the individual software product level that promote interoperability, quality improvement, porting, testing, and deployment. The U.S. Exascale Computing Project (ECP) developed its HPC software ecosystem using a three-pronged approach. First, the ECP adopted and invested in Spack, a package manager designed to handle complex HPC package dependencies. Second, the ECP created the Extreme Scale Scientific Software Stack, an effort that supports developing, deploying, and running scientific applications on HPC platforms. Third, the ECP supported software product communities, or software development kits, to develop and promote best practices, improve software interoperability, and other collaborative efforts. This article describes ECP contributions to HPC software ecosystem challenges.
Scientific processes rely on software as an important tool for data acquisition, analysis, and discovery. Over the years, sustainable software development practices have made progress in being considered as an integral component of research. However, management of computation-based scientific studies is often left to individual researchers who design their computational experiments based on personal preferences and the nature of the study. Here, we believe that the quality, efficiency, and reproducibility of computation-based scientific research can be improved by explicitly creating an execution environment that allows researchers to provide a clear record of traceability. This is particularly relevant to complex computational studies in high-performance computing (HPC) environments. In this article, we review the documentation required to maintain a comprehensive record of HPC computational experiments for reproducibility. We also provide an overview of tools and practices that we have developed to perform such studies around Flash-X, a multiphysics scientific software.
Here, this article introduces the new Research Software Engineering (RSEng) department at Computing in Science & Engineering. Through a conversation with the department coeditors, we highlight why RSEng matters, how it differs from industrial software engineering, what it means to be an RSE, and the scholarly and practical questions that lie ahead. Along the way, we draw on emerging literature, case studies, and community perspectives to frame the profession and practice of RSEng within computational science and engineering.
The rapid evolution of Software as a Service (SaaS), cloud computing, and artificial intelligence (AI) is transforming the electric utility industry, reshaping operations, customer engagement, and financial models. This webinar introduced how utilities can deploy advanced software solutions and AI-driven analytics to improve grid efficiency, optimize asset management, and accurately forecast demand.
Computational modeling and simulation have become indispensable scientific tools in virtually all areas of chemical, biomolecular, and materials systems research. Computation can provide unique and detailed atomic level information that is difficult or impossible to obtain through analytical theories and experimental investigations. In addition, recent advances in micro-electronics have resulted in computer architectures with unprecedented computational capabilities, from the largest supercomputers to common desktop computers. In conclusion, combined with the development of new computational domain science methodologies and novel programming models and techniques, this has resulted in modeling and simulation resources capable of providing results at or better than experimental chemical accuracy and for systems in increasingly realistic chemical environments.
The Adaptive Computing (AC) software stack supports goal-based computing, for which a simulation workload is created on the fly adapting to the results of calculations. Application-specific code defines an objective, which may be to solve an optimization problem or to train a surrogate model with minimal uncertainty. Then, the AC driver decides where in the design parameter space to run simulations to best achieve that objective. This process is iterative and online; as new data is returned from simulations, the AC driver chooses new simulations to run. The AC driver can strategically run simulations on distributed hardware resources (including high performance computing machines, cloud resources, and edge devices) to maximize throughput and obey resource constraints.