DMTN-199: Rubin Observatory Data Security Standards Implementation
In this document we describe a set of measures that we plan to take, in order to secure the data taken at Rubin Observatory to the standards set by the US funding agencies.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
In this document we describe a set of measures that we plan to take, in order to secure the data taken at Rubin Observatory to the standards set by the US funding agencies.
Brick (https://brickschema.org/) is a unified metadata schema to address the problem of building data standardization. Creating Brick models for building datasets means that the contents of the datasets are semantically described using the standard terms defined in the Brick ontology, and it will enable the benefits of data standardization, without having to recollect or reorganize the data. The challenge is that building brick models for building datasets leads to repeated manual trial and error processes, which can be time-consuming. VizBrick is a tool with a graphic/Web-based user interface that can assist users to create Brick models visually and interactively without having to understand the Resource Description Framework (RDF) syntax. VizBrick contains a web server that renders VizBrick web interface pages for browsers. The web server utilizes software components that (1) provide Brick ontology entity mapping to data column suggestions to users so that they can efficiently create their model; (2) provide keyword/Metadata-based search capability for easy find of relevant brick concepts and relations to their data columns
Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.
Operations of multi-reservoir systems are challenged in-part by the interplay of complex physical processes functioning within the watershed. The employment of intelligent systems can be of aid by linking environmental sensing, information technology, data analytics, simulation and decision support to achieve a data-to-decision flow of information. A further challenge is that watershed resources are managed for multiple purposes requiring some level of coordination among numerous resource managers, asset operators and users. System intelligence in this context relies on shared community platforms (data portals, community models), and coordinated communication between decision makers. Opportunities to enrich watershed intelligence has been the subject of a roadmapping exercise for the Department of Energy’s Water Power Technologies Office which has relied on broad stakeholder engagement. Initial phases of engagement involved personal interviews and a series of virtual group meetings, which focused on identifying opportunities to improve the intelligence of the physical infrastructure within our watersheds—examples of feedback include improved sensing of snowpack and runoff, data standards for facilitated data sharing, and better forecasting tools. The latter phase of engagement involved the conduct of a case study in the Upper Colorado River basin where key stakeholders were interviewed to map how their decisions are informed by intelligence from other basin stakeholders. Our presentation will highlight the interdisciplinary flow of information in complex watershed systems and identify physical and institutional opportunities toward the strategic operation of water infrastructure.
Since 2009, the U.S. Department of Energy (DOE) Office of Nuclear Energy's Nuclear Energy University Program (NEUP) has been at the forefront of nuclear research, specifically concentrating on advancing high-temperature gas-cooled reactor (HTGR) technologies. By Fiscal Year 2024, NEUP has authorized 36 projects dedicated to HTGR research, each contributing significantly to the enhancement of our understanding of this technology. The outcomes of these diverse projects have been disseminated through final NEUP reports, peer-reviewed journal articles, and presentations at academic conferences, forming a comprehensive tapestry of knowledge. Despite the substantial value of these findings, their dissemination has been fragmented, posing challenges for accessibility to researchers and policymakers and leading to underutilization of DOE investments. Recognizing this critical gap and its potential consequences for the future of nuclear research, the Advanced Reactor Technologies (ART) Gas-Cooled Reactor (GCR) program conducted an extensive survey of completed and ongoing HTGR NEUP projects. This survey enabled the compilation of crucial data, resulting in the development of a specialized public-access database tailored for computational fluid dynamics and system code validation, specifically designed for HTGR applications. However, the data collection process revealed a significant challenge in central data organization due to individual researchers from different institutes employing varying logics and preferences for recording and documenting experimental data. Consequently, an urgent need has been identified to establish a standardized reporting format for HTGR experimental projects. Addressing this issue is essential for enhancing collaboration, maximizing the impact of DOE investments, and ensuring the seamless advancement of HTGR technologies in nuclear research.
Since 2009, the U.S. Department of Energy (DOE) Office of Nuclear Energy's Nuclear Energy University Program (NEUP) has been at the forefront of nuclear research, specifically concentrating on advancing high-temperature gas-cooled reactor (HTGR) technologies. By Fiscal Year 2023, NEUP has authorized 35 projects dedicated to HTGR research, each contributing significantly to the enhancement of our understanding of this technology. The outcomes of these diverse projects have been disseminated through final NEUP reports, peer-reviewed journal articles, and presentations at academic conferences, forming a comprehensive tapestry of knowledge. Despite the substantial value of these findings, their dissemination has been fragmented, posing challenges for accessibility to researchers and policymakers and leading to underutilization of DOE investments. Recognizing this critical gap and its potential consequences for the future of nuclear research, the Advanced Reactor Technologies (ART) Gas-Cooled Reactor (GCR) program conducted an extensive survey of completed and ongoing HTGR NEUP projects. This survey enabled the compilation of crucial data, resulting in the development of a specialized public-access database tailored for computational fluid dynamics and system code validation, specifically designed for HTGR applications. However, the data collection process revealed a significant challenge in central data organization due to individual researchers from different institutes employing varying logics and preferences for recording and documenting experimental data. Consequently, an urgent need has been identified to establish a standardized reporting format for HTGR experimental projects. Addressing this issue is essential for enhancing collaboration, maximizing the impact of DOE investments, and ensuring the seamless advancement of HTGR technologies in nuclear research.
This work aims to improve the ability of particle accelerator researchers to develop high-performance accelerator cavity designs by creating an overall multiphysics framework that integrates and couples existing application codes. This framework will allow accelerator researchers to build multiphysics models that will optimize cavity design, improve understanding of whole-device performance, and reduce the development and fabrication costs of accelerator research. We utilize the open-source VizSchema data standard as an intermediate data structure interface layer to standardize interfaces between individual application codes. VizScema is extensively documented online, and plugins for VizSchema are available for popular visualization packages, including VisIt and ParaView. Currently, the work focuses on coupling the EM field solver COMSOL and the electron gun code MICHELLE to allow COMSOL field-solve results to be seamlessly used by MICHELLE for particle-solve. Later work will extend this integration to include other fields, particles, and thermodynamics simulation codes.
The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.
Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.
This dataset contains processed, standardized data from the UND scanning Doppler lidar at WFIP3's BARG site, consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.
This dataset contains processed, standardized data from the ANL scanning Doppler lidar, consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.
This dataset contains processed, standardized data from the UND scanning Doppler lidar at WFIP3's BARG site, consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.
This dataset contains standardized data from the PNNL scanning Doppler lidar (S/N 184), consisting of range- and time-resolved measurements of radial velocity, attenuated backscatter, intensity, and spectral width. We note that the beam azimuth angles are NOT referenced to true north.
Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.
The traditional approach to planning the distribution grid has focused on reliability in the context of gradual and reasonably predictable load growth. Forecasts of load growth, combined with asset management practices, were used by system planners to identify upgrades to the system to maintain or improve reliability. The decisions, typically based within load flow analysis tools, included considerations about contingency scenarios and corporate forecasts (i.e., top-down predictions at a summary level of what would happen in a particular area that could impact load growth and behavior). As a result, today, this traditional approach no longer fits all purposes.
Not provided.
At the IEEE/ACM International Conference for High-Performance Computing, Networking, Storage, and Analysis (SC23), held in Denver, experts discussed the convergence of high-performance computing and cloud computing. Experts explored how this integration could address current scientific computing limitations, enhance computational capabilities, and foster global collaboration while focusing on economic, security, technical, and community challenges and opportunities.
City Buildings, Energy, and Sustainability (CityBES) is a web-based data and computing platform, focusing on energy modeling and analysis of a city's building stock to support district or city-scale building energy efficiency programs. CityBES uses an international open data standard, CityGML, to represent and exchange 3D city models. CityBES employs EnergyPlus to simulate building energy use and savings from energy efficient retrofits. Other CityBES features include energy benchmarking, district heating and cooling system modeling, rooftop PV analysis, building performance visualization, heat resilience modeling, as well as urban scale mapping of microclimate and heat vulnerability at census tract level. Different from other tools, CityBES uses integrated open and standard 3D city building data and models each individual building using EnergyPlus. CityBES can be used by urban planners, city energy managers, building owners, utilities, energy consultants and researchers.