Enabling rapid development of highly interactive analysis tools for accelerator operations and experiments at DARHT through web services [Slides]
The penetrating radiography provided by DARHT is a key capability in executing a core mission of LANL.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
The penetrating radiography provided by DARHT is a key capability in executing a core mission of LANL.
Explore the source record for details and available documents.
Long-term, high-resolution, regional wave hindcast datasets were generated using unstructured-grid Simulating WAves Nearshore (SWAN) models for the U.S. coastal waters to support nearshore wave energy development in the U.S. including those bordering U.S. territorial islands. The model domains resolved the entire U.S. exclusive economic zones, with a spatial resolution of approximately 200 m nearshore. The regional SWAN models were driven by the global WAVEWATCH III® model outputs and run for a 42-year period from 1979 to 2020. Extensive model validations were performed using buoy observations and altimeter data. Regional resource characterization was performed based on hindcast data points at 2 km from shore and along the 100 m isobath. Aggregations of wave resource parameters were produced, and spatial and seasonal variations were analyzed for all the regions. Wave resource metrics recommended by international standards, including a 3-hour time series of six resource parameters, hourly frequency- and directionally resolved wave spectra at selected “virtual buoy” locations, and average-annual values of omni-directional wave power, significant wave height, and energy period are publicly disseminated through an Amazon Web Service and a Marine Energy Atlas web application tool to facilitate wave energy research and a wide range of coastal ocean applications.
The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.
Transitioning from an onsite Maintenance & Diagnostics Center to cloud-based services offers many new opportunities with computing power and storage, but also new challenges in terms of networking and security. This report will cover everything required for that transition including data processing and uploading to cloud services, feature selection, model creation, and result visualization for decision making. Although there are several other cloud-based services (e.g. Amazon Web Services and Google Cloud), this report explores Microsoft Azure to simplify nomenclature and maintain a consistent focus. Many of the services offered by Microsoft Azure are also available in the other cloud-based services, and their differences have been recorded in other literature. The Azure services most important to a nuclear power plant including networking & security, storage & databases, and Artificial Intelligence (AI) are reviewed here. Networking covers all aspects related to communication to Azure resources including security, privacy, and redundancy. Storage & databases includes data storage, upgrading, patching, backups, and monitoring. The AI services allows the user access to the machine learning (ML) techniques developed with Azure including automated ML, anomaly detection, computer vision, and natural language processing. This report summaries the features, capabilities, and challenges when using cloud-based services in a user-friendly manner.
The ATLAS experiment at CERN is one of the largest scientific machines built to date and will have ever growing computing needs as the Large Hadron Collider collects an increasingly larger volume of data over the next 20 years. ATLAS is conducting R&D projects on Amazon Web Services and Google Cloud as complementary resources for distributed computing, focusing on some of the key features of commercial clouds: lightweight operation, elasticity and availability of multiple chip architectures. The proof of concept phases have concluded with the cloud-native, vendoragnostic integration with the experiment’s data and workload management frameworks. Google Cloud has been used to evaluate elastic batch computing, ramping up ephemeral clusters of up to O(100k) cores to process tasks requiring quick turnaround. Amazon Web Services has been exploited for the successful physics validation of the Athena simulation software on ARM processors. We have also set up an interactive facility for physics analysis allowing endusers to spin up private, on-demand clusters for parallel computing with up to 4 000 cores, or run GPU enabled notebooks and jobs for machine learning applications. The success of the proof of concept phases has led to the extension of the Google Cloud project, where ATLAS will study the total cost of ownership of a production cloud site during 15 months with 10k cores on average, fully integrated with distributed grid computing resources and continue the R&D projects.
The proliferation of low-cost sensors and industrial data solutions has continued to push the frontier of manufacturing technology. Machine learning and other advanced statistical techniques stand to provide tremendous advantages in production capabilities, optimization, monitoring, and efficiency. The tremendous volume of data gathered continues to grow, and the methods for storing the data are critical underpinnings for advancing manufacturing technology. This work aims to investigate the ramifications and design tradeoffs within a decoupled architecture of two prominent database management systems (DBMS): sql and NoSQL. A representative comparison is carried out with Amazon Web Services (AWS) DynamoDB and AWS Aurora MySQL. The technologies and accompanying design constraints are investigated, and a side-by-side comparison is carried out through high-fidelity industrial data simulated load tests using metrics from a major US manufacturer. The results support the use of simulated client load testing for comparing the latency of database management systems as a system scales up from the prototype stage into production. As a result of complex query support, MySQL is favored for higher-order insights, while NoSQL can reduce system latency for known access patterns at the expense of integrated query flexibility. Here, by reviewing this work, a manufacturer can observe that the use of high-fidelity load testing can reveal tradeoffs in IoTfM write/ingestion performance in terms of latency that are not observable through prototype-scale testing of commercially available cloud DB solutions.
Federated learning enables multiple data owners to collaboratively train robust machine learning models without transferring large or sensitive local datasets by only sharing the parameters of the locally trained models. Here, in this article, we elaborate on the design of our Advanced Privacy-Preserving Federated Learning (APPFL) framework, which streamlines end-to-end secure and reliable federated learning experiments across cloud computing facilities and high-performance computing resources by leveraging Globus Compute, a distributed function as a service platform, and Amazon Web Services. We further demonstrate the use case of APPFL in fine-tuning an LLaMA 2 7B model using several cloud resources and supercomputers.
This report continues and extends the development of the Energy Services Interface (ESI) that recently has been led by the U.S. Department of Energy’s Grid Modernization Laboratory Consortium. The ESI is intended to facilitate the coordination of multiple, flexible energy resources that work in tandem to satisfy grid objectives by using performance attributes encoded within an ESI Service Template. The ESI relies on a pair of interfaces, representing the service requestor (the consumer of a service) and a service provider (who manages its resources per the agreed performance attributes) to achieve what is commonly known as a “grid service”. Due to its performance-driven approach, the ESI is hypothesized to be able to satisfy all common grid needs via only six common ESI service types. This report first reviews the status of ESI development, including its fundamental tenets. The report makes three important contributions to ESI development: First, it recognizes the similarity between service level agreements and the contract-like agreements that would be needed between energy service requestors and providers and recommends that ESI service agreements be modeled after the web services agreement specification. Second, whereas prior development efforts had focused on a flat, five-stage lifecycle, this report asserts that the ESI should have three behavioral layers in its architecture—the discovery, agreement, and service layers. Finally, this report offers concrete requirements that should be useful toward the development of the service requestors’ and service providers’ respective communications interfaces.
Buildings are active participants in increasingly complex energy systems. Building Energy Modeling (BEM) has a key role to play in planning and de-risking an equitable energy transition, with BEM-backed "virtual buildings" critical path for diverse applications that include workforce training tools, Hardware-in-the-Loop (HIL) experimentation to study equipment performance under a range of conditions, Control-Hardware-in-the-Loop (CHIL) experimentation to de-risk commercial control implementations at equipment through grid orchestration levels, and integration of dynamic load profiles into grid modeling tools for energy system experimentation at the urban scale. Modeling requirements vary across these applications, but many software engineering tasks do not. The Alfalfa Virtual Building Service (AVBS, see https://github.com/NREL/alfalfa/wiki) is an open-source web service that solves these common tasks robustly in one place, providing a foundational platform for power users to bootstrap their own applications. AVBS abstracts the specifics of runtime interaction with OpenStudio, Modelica, and Spawn of EnergyPlus models behind a unified REST API. Additionally, AVBS provides resources for cloud deployment and scaling to 100s of parallel simulations, a growing library of modular Operational Technology (OT) integrations for emulation of real-world interfaces, and scripts to automate the population of communities of virtual buildings from URBANopt, ResStock and ComStock.
The operation of the neutron facility relies heavily on beamline scientists. Some experiments can take one or two days with experts making decisions along the way. Leveraging the computing power of HPC platforms and AI advances in image analyses, here we demonstrate an autonomous workflow for the single-crystal neutron diffraction experiments. The workflow consists of three components: an inference service that provides real-time AI segmentation on the image stream from the experiments conducted at the neutron facility, a continuous integration service that launches distributed training jobs on Summit to update the AI model on newly collected images, and a frontend web service to display the AI tagged images to the expert. Ultimately, the feedback can be directly fed to the equipment at the edge in deciding the next-step experiment without requiring an expert in the loop. With the analyses of the requirements and benchmarks of the performance for each component, this effort serves as the first step toward an autonomous workflow for real-time experiment steering at ORNL neutron facilities.
Abstract We present a scalable, cloud-based science platform solution designed to enable next-to-the-data analyses of terabyte-scale astronomical tabular data sets. The presented platform is built on Amazon Web Services (over Kubernetes and S3 abstraction layers), utilizes Apache Spark and the Astronomy eXtensions for Spark for parallel data analysis and manipulation, and provides the familiar JupyterHub web-accessible front end for user access. We outline the architecture of the analysis platform, provide implementation details and rationale for (and against) technology choices, verify scalability through strong and weak scaling tests, and demonstrate usability through an example science analysis of data from the Zwicky Transient Facility’s 1Bn+ light-curve catalog. Furthermore, we show how this system enables an end user to iteratively build analyses (in Python) that transparently scale processing with no need for end-user interaction. The system is designed to be deployable by astronomers with moderate cloud engineering knowledge, or (ideally) IT groups. Over the past 3 yr, it has been utilized to build science platforms for the DiRAC Institute, the ZTF partnership, the LSST Solar System Science Collaboration, and the LSST Interdisciplinary Network for Collaboration and Computing, as well as for numerous short-term events (with over 100 simultaneous users). In a live demo instance, the deployment scripts, source code, and cost calculators are accessible. 4 4 http://hub.astronomycommons.org/
The PV Operations and Maintenance (O&M) service industry lacks an affordable, well-documented, intuitive PV modeling and analytics tool to calculate modeled performance from actual data from multiple data acquisition systems (DAS). We envision a performance modeling and analytics platform built on open-source, extensible, community-maintained code. The key innovation is the community-driven development of pvlib python delivered through a lightweight web service to provide configurable, consistent and reproducible PV modeling for O&M providers.
Machine learning enjoys widespread success in High Energy Physics (HEP) analyses at LHC. However the ambitious HL-LHC program will require much more computing resources in the next two decades. Quantum computing may offer speed-up for HEP physics analyses at HL-LHC, and can be a new computational paradigm for big data analyses in High Energy Physics.We have successfully employed three methods (1) Variational Quantum Classifier (VQC) method, (2) Quantum Support Vector Machine Kernel (QSVM-kernel) method and (3) Quantum Neural Network (QNN) method for two LHC flagship analyses: ttH (Higgs production in association with two top quarks) and H->mumu (Higgs decay to two muons, the second generation fermions). We shall address the progressive improvements in performance from method (1) to method (3).We will present our experiences and results of a study on LHC High Energy Physics data analyses with IBM Quantum Simulator and Quantum Hardware (using IBM Qiskit framework), Google Quantum Simulator (using Google Cirq framework), and Amazon Quantum Simulator (using Amazon Braket cloud service). The work is in the context of a Qubit platform (a gate-model quantum computer). Taking into account the present limitation of hardware access, different quantum machine learning methods are studied on simulators and the results are compared with classical machine learning methods (BDT, classical Support Vector Machine and classical Neural Network). Furthermore, we do apply quantum machine learning on IBM quantum hardware to compare performance between quantum simulator and quantum hardware. The work is performed by an international and interdisciplinary collaboration with the Department of Physics and Department of Computer Sciences of University of Wisconsin, CERN Quantum Technology Initiative, IBM Research Zurich, IBM T.J. Watson Research Center, Fermilab Quantum Institute, BNL Computational Science Initiative, State University of New York at Stony Brook, and Quantum Computing and AI Research of Amazon Web Services. This work pioneers a close collaboration of academic institutions with industrial corporations in the High Energy Physics analyses effort. Though the size of event samples in future HL-LHC physics and the limited number of qubits pose some challenges to the Quantum Machine learning studies for High Energy Physics, more advanced quantum computers with larger number of qubits, reduced noise and improved running time (as envisioned by IBM and Google) may outperform classical machine learning in both classification power and in speed.Although the era of efficient quantum computing may still be years away, we have made promising progress and obtained preliminary results in applying quantum machine learning to High Energy Physics. A PROOF OF PRINCIPLE.
This document describes the data products and processing services to be delivered by the NSF-DOE Vera C. Rubin Observatory whilst performing the Legacy Survey of Space and Time (LSST). LSST will deliver three levels of data products and services. Prompt data products are computed and released within 24 hours of observation, and include images, difference images, catalogs of sources and objects detected in difference images, and catalogs of Solar System objects. Their primary purpose is to enable rapid follow-up of time-domain events. Data Release data products are computed during annual processing campaigns, and include well-calibrated single-epoch images, deep coadds, and catalogs of objects, sources, and forced sources, enabling static sky and precision time-domain science. The Science Platform will allow for the creation of User Generated data products and will enable science cases that greatly benefit from co-location of user processing and/or data within the Rubin Observatory Data Access Center. LSST will also devote 10% of observing time to programs with special cadence. Their data products will be created using the same software and hardware as Prompt and Data Release products. All data products will be made available using user-friendly databases and web services.
With an increased dataset obtained during the Run 3 of the LHC at CERN and the even larger expected increase of the dataset by more than one order of magnitude for the HL-LHC, the ATLAS experiment is reaching the limits of the current data processing model in terms of traditional CPU resources based on x86_64 architectures and an extensive program for software upgrades towards the HL-LHC has been set up. The ARM architecture is becoming a competitive and energy efficient alternative. Some surveys indicate its increased presence in HPCs and commercial clouds, and some WLCG sites have expressed their interest. Chip makers are also developing their next generation solutions on ARM architectures, sometimes combining ARM and GPU processors in the same chip. Consequently it is important that the ATLAS software embraces the change and is able to successfully exploit this architecture. We report on the successful porting to ARM of the Athena software framework, which is used by ATLAS for both online and offline computing operations. Furthermore we report on the successful validation of simulation workflows running on ARM resources. For this we have set up an ATLAS Grid site using ARM compatible middleware and containers on Amazon Web Services (AWS) ARM resources. The ARM version of Athena is fully integrated in the regular software build system and distributed in the same way as other software releases. In addition, the workflows have been integrated into the HEPscore benchmark suite which is the planned WLCG wide replacement of the HepSpec06 benchmark used for Grid site pledges. In the overall porting process we have used resources on AWS, Google Cloud Platform (GCP) and CERN. A performance comparison of different architectures and resources will be discussed.
Abstract Hydroelectric power (hydropower) is unique in that it can function as both a conventional source of electricity and as backup storage (pumped hydroelectric storage and large reservoir storage) for providing energy in times of high demand on the grid (S. Rehman, L M Al-Hadhrami, and M M Alam), (2015 Renewable and Sustainable Energy Reviews , 44 , 586–98). This study examines the impact of hydropower on system electricity price and price volatility in the region served by the New England Independent System Operator (ISONE) from 2014-2020 (ISONE, ISO New England Web Services API v1.1 .” https://webservices.iso-ne.com/docs/v1.1/ , 2021. Accessed: 2021-01-10). We perform a robust holistic analysis of the mean and quantile effects, as well as the marginal contributing effects of hydropower in the presence of solar and wind resources. First, the price data is adjusted for deterministic temporal trends, correcting for seasonal, weekend, and diurnal effects that may obscure actual representative trends in the data. Using multiple linear regression and quantile regression, we observe that hydropower contributes to a reduction in the system electricity price and price volatility. While hydropower has a weak impact on decreasing price and volatility at the mean, it has greater impact at extreme quantiles (>70th percentile). At these higher percentiles, we find that hydropower provides a stabilizing effect on price volatility in the presence of volatile resources such as wind. We conclude with a discussion of the observed relationship between hydropower and system electricity price and volatility.
Graphs are a natural and fundamental representation to describe entities, relationships, activities, and evolution of complex systems. Many domains such as communication, citation, procurement, biology, social media, and transportation can be modeled as a set of entities and their relationships. Resource Description Framework (RDF) and Labeled Property Graph (LPG) are two of the most used data models to encode information in a graph. Both models are similar in terms of using basic graph elements such as nodes and edges but differ in terms of the modeling approach, expressibility, serialization, and target applications. RDF is a flexible data exchange model for expressing information about entities but it tends to a have high memory footprint and inefficient storage, which does not make it a natural choice to perform scalable graph analytics. In contrast, LPG has gained traction as a reliable model to perform scalable graph analytic tasks such as sub-graph matching, network alignment, and real-time knowledge graph query. It provides efficient storage, fast traversal, and flexibility to model various real-world domains. At the same time, the LPG lacks the support of a formal knowledge representation such as an ontology to provide automated knowledge inference. We propose Semantic Property Graph (SPG) as a logical projection of reified RDF into the LPG model. SPG continues to use RDF ontology to define the type hierarchy of the projected graph and validate it against a given ontology. We present a framework to convert reified RDF graphs into SPG using two different computing environments. We also present cloud-based graph migration capabilities using Amazon Web Services.