Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “open data format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Dataset for scientific paper "Simulated plant‑mediated oxygen input has strong impacts on fine‑scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands", a modeling study based on field observation at the tidal salt marshes of the Parker River Estuary, Massachusetts, United States

This dataset is the raw and processed data for the paper "Simulated plant ‑ mediated oxygen input has strong impacts on fine ‑ scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands". This study investigated how plant-mediated oxygen input affects subsurface biogeochemical reactions of organic carbon degradation and the resulting methane emissions of coastal wetlands by model simulation. We used the subsurface geochemical simulator PFLOTRAN for the modeling, which produced the simulated changes in porewater chemical substances and methane emissions over 10 days under different scenarios of plant-mediated oxygen input.Specifically, this dataset contains: 1) the input files for PFLOTRAN of all simulation runs conducted in this study. Those files are with an extension of ".in", containing information of the biogeochemical reaction network (stoichiometry, reaction rate, Monod constants, etc), fluid flow rate and oxygen concentration in the fluid which together simulated the plant-mediated oxygen input, the configuration of artificial reactions that simulated the methane fluxes, etc. The PFLOTRAN input files are text files, which can be opened by NotePad, but running these input files will require proper installation of PFLOTRAN (instruction: https://documentation.pflotran.org/user_guide/how_to/installation/installation.html). 2) the raw and processed model output from PFLOTRAN of all simulation runs, and 3) the python scripts used to process the raw model output, including random allocation of root cells, converting raw data into organized formats, calculating the methane fluxes based on the model output, data visualization, etc. The raw and processed model output from PFLOTRAN are in .spydata format, which can be viewed with Python. and 3) the python scripts for data processing and analysis are programming scripts, which can be opened with Python.This modeling work, in particular the model parameterization of root density and initial conditions of porewater concentrations of biogeochemical substances, was based on field measurements at the salt marsh of the Upper Parker River Estuary, Massachusetts, United States.

54 ENVIRONMENTAL SCIENCES↗

Conquering Data Chaos: Research Data Management with Kubernetes

Managing massive volumes of data and effectively making it accessible to researchers poses significant challenges and is a barrier to scientific discovery. In many cases, critical data is locked up in unwieldy file formats or one-off databases and is too large to effectively process on a single machine. This talk explores the role of Kubernetes, an open-source container orchestration platform, in addressing research data management challenges. I will discuss how we are using a set of publicly available open-source and home-grown tools in the National Renewable Energy Lab (NREL) Data, Analysis, and Visualization (DAV) group to help researchers overcome data-related bottlenecks. The talk will begin by providing an overview of the data challenges faced in research data management, including data storage, processing, and analysis. I will highlight Kubernetes' ability to handle large-scale data by leveraging containerization and distributed computing, including distributed storage. Kubernetes allows researchers to encapsulate data processing infrastructure and workflows into portable containers, enabling reproducibility and ease of deployment. Kubernetes can then schedule and manage the resource allocation of these containers to enable efficient utilization of limited computing resources, leading to more efficient data processing and analysis. I will discuss some limitations of traditional, siloed approaches to dealing with data and emphasize the need for solutions which foster collaboration. I will highlight how we are using Kubernetes at NREL to facilitate data sharing and cooperation among research teams. Kubernetes' flexible architecture enables the deployment of shared computing environments, such as Apache Superset, where researchers can seamlessly access and analyze shared datasets. Providing the ability to have one research team easily consume data generated by another, utilizing Kubernetes' as a central data platform, is one of the major wins we've encountered by adopting the platform. Finally, I will showcase real-world use cases from NREL where we have used Kubernetes to solve some persistent data challenges involving large volumes of sensor and monitoring data. I will discuss the challenges we encountered when creating our cluster and making it available as a production-ready resource. I will also discuss the specific suite of tools, including Postgres and Apache Druid for columnar and timeseries data, and Redpanda Kafka for streaming data we have deployed in our infrastructure, and the process that went into the selection of these tools.

collaborative environment↗

Spurious solar-wind effects on acceleration noise in LISA Pathfinder

Spurious solar-wind effects are a potential noise source in future Laser Interferometer Space Antenna (LISA) measurements. One noise coupling mechanism is constrained by estimating solar-wind effects on acceleration noise in LISA Pathfinder (LPF). While LISA is designed for drag-free differential measurement, predicting the realistic impact both bounds the operational environment and assesses whether LISA could provide serendipitous space-weather observations. Data from NASA's Advanced Composition Explorer (ACE), situated at the L1 Lagrange point, serves as a reliable source of solar-wind data. The data sets are compared over the 114 d time period from 1 March 2016 to 23 June 2016. This period gives the longest readily-available open data set, without interference from other commissioning activities. To evaluate space weather effects, the data from both satellites are formatted, gap-filled/interpolated, and fast-Fourier transformed for amplitude spectral density and coherence comparisons. Solar wind effects are not seen in a coherence plot between LPF and ACE; modest coherence in the planned LISA observational frequency band can be attributed to chance. This result indicates that measurable correlation due to solar-wind acceleration noise over 3 month timescales will be a negligible noise source. LISA is unlikely to inform solar wind measurements routinely. Another source of noise from the Sun, solar radiation pressure, is estimated to impart greater acceleration noise, but has yet to be analyzed.

79 ASTRONOMY AND ASTROPHYSICS↗

The Foundational Industrial Energy Dataset (FIED): Open-Source Data on Industrial Facilities

The state of data on industrial energy use has co-evolved over several decades with the demands of industrial energy analysis. The most recent development - analysis in support of decarbonizing the industrial sector - has changed the characteristics of industrial data that are useful for analysts and model developers. Although data and its collection processes may be cast from a conventional viewpoint as objective and free from the influence of social dynamics, this provides an incomplete picture of not only the processes by which information is generated, but also the limitations and opportunities of data to be useful for analysis. The foundational industry energy data set (FIED) is a result of the confluence of trends in open data and the demand for higher resolution industrial energy analysis. The general approach to compiling the FIED involves accessing, filtering, and formatting data published by federal organizations on the Internet for public use. Unlike most industrial energy datasets, which are published by the U.S. Energy Information Administration (EIA), the FIED relies on core datasets from the U.S. Environmental Protection Agency (EPA). The FIED addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization - including estimates of energy use, greenhouse gas emissions, and design capacities - for facilities that are identified by latitude and longitude. This enables local-level analysis of existing combustion equipment, as well as regional comparisons with traditional industrial energy data estimates. The report summarizes the general logic behind compiling the FIED. The FIED itself and its Python code are available from OpenEI and GitHub, respectively.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Dissolution zone model of the oxide structure in additively manufactured dispersion-strengthened alloys

The structural evolution of oxides in dispersion-strengthened superalloys during laser-powder bed fusion is considered in detail. Alloy chemistry and process parameter effects on oxide structure are assessed through a parameter study on the model alloy Ni-20Cr, doped with varying concentrations of Y 2 O 3 and Al. Small angle neutron scattering measurements of the dispersoid size distribution show the dispersoid size increases with higher laser power, slower scan speed, and increasing Y 2 O 3 and Al content. Complementary electron microscopy measurements reveal reactions between Y 2 O 3 and Al, even in nanoscale dispersoids, and the presence of micron-scale oxide slag inclusions in select specimens. A scaling analysis of mass and momentum transport within the melt pool, presented here, establishes that diffusional structural evolution mechanisms dominate for nanoscale dispersoids, while fluid forces and advection become significant for larger slag inclusions. These findings are developed into a theory of dispersoid structural evolution, integrating quantitative models of diffusional processes – dispersoid dissolution, nucleation, growth, coarsening – with a reduced order model of time-temperature trajectories of fluid parcels within the melt pool. Calculations of the dispersoid size in single-pass melting reveal a zone in the center of the melt track in which the oxide feedstock fully dissolves. Within this zone the final Y 2 O 3 size is independent of feedstock size and determined by nucleation and growth kinetics. If the dissolution zones of adjacent melt tracks overlap sufficiently with each other to dissolve large oxides, formed during printing or present in the powder feedstock, then the dispersoid structure throughout the build volume is homogeneous and matches that from a single pass within the dissolution zone. Gaps between adjacent dissolution zones result in oxide accumulation into larger slag inclusions. Predictions of final dispersoid size and slag formation using this dissolution zone model match the present experimental data and explain process-structure linkages speculated in the open literature.

36 MATERIALS SCIENCE↗

A Processing and Analytics System for Microscopy Data Workflows: The Pycroscopy Ecosystem of Packages

Major advancements in fields as diverse as biology and quantum computing have relied on a multitude of microscopy techniques. Despite the considerable proliferation of these instruments, significant bottlenecks remain in terms of processing, analysis, storage, and retrieval of the acquired datasets. Aside from lack of file standards, individual domain-specific analysis packages are often disjoint from the underlying datasets, and thus keeping track of analysis and processing steps remains tedious for the end-user, hampering reproducibility. Here, in this study, the pycroscopy ecosystem of packages is introduced, an open-source python-based ecosystem underpinned by a common data model. The data model, termed the N-dimensional spectral imaging data format, is realized in pycroscopy's sidpy package. This package is built on top of dask arrays, thus leveraging dask array attributes, but expanding them to accelerate microscopy relevant analysis and visualization. Several examples of the use of the pycroscopy ecosystem to create workflows for data ingestion and analysis of scanning transmission electron microscopy (STEM) and scanning probe microscopy data are shown. Adoption of such standardized routines will be critical to usher in the next generation of autonomous instruments where processing, computation, and meta-data storage will be critical to overall experimental operations.

97 MATHEMATICS AND COMPUTING↗

Surface Complexation/Ion Exchange Hybrid Model for Radionuclide Sorption to Clay Minerals (M4SF-23LL010301062)

This progress report (Level 4 Milestone Number M4SF-23LL010301062) summarizes research conducted at Lawrence Livermore National Laboratory (LLNL) within the Argillite International Collaborations Activity Number SF-23LL01030106. The activity is focused on our long-term commitment to engaging our partners in international nuclear waste repository research. The focus of this milestone is the establishment of international collaborations for surface complexation modeling and the associated impacts of unlocking larger, community-based datasets. More specifically, we are developing a database framework for Spent Fuel and Waste and Science Technology (SFWST) that is aligned with the Helmholtz Zentrum Dresden Rossendorf (HZDR) sorption database development group in support of the database needs of the SFWST program. In our FY22 effort, we described a detailed analysis of U(VI) sorption to quartz through both traditional surface complexation modeling and through a hybrid ML framework. In FY23, effort was placed on publication of these results and expansion of the LLNL surface complexation and ion exchange database (L-SCIE) in order to assess mineral-based radionuclide retardation under a wider variety of geochemical conditions (e.g., ionic strength, varying electrolyte compositions). Efforts were initiated to expand L-SCIE to include radionuclide surface complexation and ion exchange to clays that are relevant to subsurface geochemical processes occurring at nuclear waste repositories. In particular, a large source of sorption data for clays resides at the Paul Scherrer Institute (PSI) (work primarily by Bradbury and Baeyens) and we initiated discussions on how to retrieve those data and apply FAIR principles to those datasets. In addition to L-SCIE development, two hybrid models that incorporate AI/ML were investigated and compared to discern the most promising approaches for accurate and precise estimations of radionuclide retardation. Key considerations for future model development include (1) the ability to reduce computational burden on determining retardation coefficients for PA and (2) the ability to quantify and predict radionuclide-mineral partitioning at a more efficient, rapid pace due to automated workflows. Upon the careful consideration of the most effective modeling approaches, we are identifying ways to implement these approaches into PA. Ultimately, the data science-based workflows will provide a major incentive for other institutions to adopt a FAIR-formatted, interoperable database. LLNL will play a key role in disseminating sorption data and acting as good data stewards by updating the database in a consistent format and assessing the quality of the newly assimilated data in an organized fashion. To this end, all data and workflows are open access and made available on the LLNL Seaborg research website (https://seaborg.llnl.gov/resources/geochemical-databases-modeling-codes).

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Crystal structures reveal catalytic and regulatory mechanisms of the dual-specificity ubiquitin/FAT10 E1 enzyme Uba6

The E1 enzyme Uba6 initiates signal transduction by activating ubiquitin and the ubiquitin-like protein FAT10 in a two-step process involving sequential catalysis of adenylation and thioester bond formation. To gain mechanistic insights into these processes, we determined the crystal structure of a human Uba6/ubiquitin complex. Two distinct architectures of the complex are observed: one in which Uba6 adopts an open conformation with the active site configured for catalysis of adenylation, and a second drastically different closed conformation in which the adenylation active site is disassembled and reconfigured for catalysis of thioester bond formation. Surprisingly, an inositol hexakisphosphate (InsP6) molecule binds to a previously unidentified allosteric site on Uba6. Our structural, biochemical, and biophysical data indicate that InsP6 allosterically inhibits Uba6 activity by altering interconversion of the open and closed conformations of Uba6 while also enhancing its stability. In addition to revealing the molecular mechanisms of catalysis by Uba6 and allosteric regulation of its activities, our structures provide a framework for developing Uba6-specific inhibitors and raise the possibility of allosteric regulation of other E1s by naturally occurring cellular metabolites.

59 BASIC BIOLOGICAL SCIENCES↗

COMPASS-FME Synoptic Sites Level 2 Sensor Data v2-1

This is the version 2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our synoptic field sites. COMPASS-FME is studying sites in two distinct regions, the Chesapeake Bay and the Western Lake Erie Basin. We established the network at seven "synoptic" (observational) sites along the Chesapeake Bay and Lake Erie coastlines, collectively generating over three million observations per month, to track and comprehend environmental changes where land and water intersect. Additionally, the two regions provide an interesting contrast of saltwater and freshwater coasts that allow us to differentiate the impacts of inundation and coastal water chemistries in two nationally important coastal systems. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. Please see v2-1 L2 Sensor Package QStart.pdf for detailed information on data package structure, temporal coverage, and versioning.

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

Generalizable Web User Interface for Scalable and Streamlined Deployment of Building Energy Management Systems in Small and Medium-Sized Commercial Buildings

Small and medium-sized commercial buildings (SMCBs) comprise 94% of US commercial buildings yet face significant barriers to implementing building energy management systems despite advances in smart device technology. Existing solutions present critical limitations: cloud-based API solutions simplify deployment but create vendor lock-in constraints; commercial integrated software solutions ensure compatibility via standardized protocols but require substantial cost and technical expertise; open-source IoT platforms offer cost-effective vendor independence but provide insufficient standardized protocol support for commercial building automation. This research presents a generalizable web user interface framework that bridges the gap between evolving smart device capabilities and lagging software infrastructure for SMCBs. The proposed system integrates VOLTTRON open-source middleware with an automated configuration converter that transforms unified specifications written in YAML, a human-readable data-serialization format, into system-specific files, streamlining manual setup processes. The vendor-agnostic architecture supports industry-standard protocols (BACnet and Modbus) and semantic building models while providing adaptive web interfaces that dynamically adjust to various building configurations. Demonstrations through simulation-based testing and a field deployment show automatic interface adaptation across heterogeneous HVAC systems and multizone monitoring. The automated configuration converter also substantially reduces labor-intensive setup.

Chung, Jihoon [ORNL] (ORCID:0000000184880815)↗

Symphony: Cosmological Zoom-in Simulation Suites over Four Decades of Host Halo Mass

Abstract We present Symphony, a compilation of 262 cosmological, cold-dark-matter-only zoom-in simulations spanning four decades of host halo mass, from 10 11 –10 15 M ⊙ . This compilation includes three existing simulation suites at the cluster and Milky Way–mass scales, and two new suites: 39 Large Magellanic Cloud-mass (10 11 M ⊙ ) and 49 strong-lens-analog (10 13 M ⊙ ) group-mass hosts. Across the entire host halo mass range, the highest-resolution regions in these simulations are resolved with a dark matter particle mass of ≈3 × 10 −7 times the host virial mass and a Plummer-equivalent gravitational softening length of ≈9 × 10 −4 times the host virial radius, on average. We measure correlations between subhalo abundance and host concentration, formation time, and maximum subhalo mass, all of which peak at the Milky Way host halo mass scale. Subhalo abundances are ≈50% higher in clusters than in lower-mass hosts at fixed sub-to-host halo mass ratios. Subhalo radial distributions are approximately self-similar as a function of host mass and are less concentrated than hosts’ underlying dark matter distributions. We compare our results to the semianalytic model Galacticus , which predicts subhalo mass functions with a higher normalization at the low-mass end and radial distributions that are slightly more concentrated than Symphony. We use UniverseMachine to model halo and subhalo star formation histories in Symphony, and we demonstrate that these predictions resolve the formation histories of the halos that host nearly all currently observable satellite galaxies in the universe. To promote open use of Symphony, data products are publicly available at http://web.stanford.edu/group/gfc/symphony .

79 ASTRONOMY AND ASTROPHYSICS↗

Tractometry of the Human Connectome Project: resources and insights

The Human Connectome Project (HCP) has become a keystone dataset in human neuroscience, with a plethora of important applications in advancing brain imaging methods and an understanding of the human brain. We focused on tractometry of HCP diffusion-weighted MRI (dMRI) data. We used an open-source software library (pyAFQ; https://yeatmanlab.github.io/pyAFQ) to perform probabilistic tractography and delineate the major white matter pathways in the HCP subjects that have a complete dMRI acquisition (n = 1,041). We used diffusion kurtosis imaging (DKI) to model white matter microstructure in each voxel of the white matter, and extracted tract profiles of DKI-derived tissue properties along the length of the tracts. We explored the empirical properties of the data: first, we assessed the heritability of DKI tissue properties using the known genetic linkage of the large number of twin pairs sampled in HCP. Second, we tested the ability of tractometry to serve as the basis for predictive models of individual characteristics (e.g., age, crystallized/fluid intelligence, reading ability, etc.), compared to local connectome features. To facilitate the exploration of the dataset we created a new web-based visualization tool and use this tool to visualize the data in the HCP tractometry dataset. Finally, we used the HCP dataset as a test-bed for a new technological innovation: the TRX file-format for representation of dMRI-based streamlines. We released the processing outputs and tract profiles as a publicly available data resource through the AWS Open Data program's Open Neurodata repository. We found heritability as high as 0.9 for DKI-based metrics in some brain pathways. We also found that tractometry extracts as much useful information about individual differences as the local connectome method. We released a new web-based visualization tool for tractometry—“Tractoscope” (https://nrdg.github.io/tractoscope). We found that the TRX files require considerably less disk space-a crucial attribute for large datasets like HCP. In addition, TRX incorporates a specification for grouping streamlines, further simplifying tractometry analysis.

59 BASIC BIOLOGICAL SCIENCES↗

The 2025 “Hacking Limnology” Workshop Series and DSOS Virtual Summit: A Half Decade of Data‐Intensive Aquatic Science

The 5th Aquatic Ecosystem MOdeling Network—Junior (AEMON-J) “Hacking Limnology” Workshop and 6th Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) convened 21–25 July 2025. As in previous years (Fig. 1; Meyer and Zwart 2020; Meyer et al. 2021b, 2021c, 2022, 2024), the virtual workshops and summit were free of charge, the content was formatted to allow for broad engagement from a globally distributed audience, and workshop materials and recordings were made available on the AEMON-J/DSOS archive (Meyer et al. 2021a). In contrast to previous years, which primarily focused on inland aquatic ecosystems, this year's workshops and summit showcased a notable plurality of ecosystem types, with workshops spanning marine, riverine, and lacustrine environments. The weeklong event brought together researchers and practitioners interested in the nexus of data science, open science, and the aquatic sciences, hosting between 47 and 65 attendees at a single time and a higher number of registrants (n = 389), who might opt to access the material asynchronously.

Meyer, Michael F. [US Geological Survey, Portland,↗

Dynamic STEM-EELS for single-atom and defect measurement during electron beam transformations

This study introduces the integration of dynamic computer vision–enabled imaging with electron energy loss spectroscopy (EELS) in scanning transmission electron microscopy (STEM). This approach involves real-time discovery and analysis of atomic structures as they form, allowing us to observe the evolution of material properties at the atomic level, capturing transient states traditional techniques often miss. Rapid object detection and action system enhances the efficiency and accuracy of STEM-EELS by autonomously identifying and targeting only areas of interest. This machine learning (ML)–based approach differs from classical ML in that it must be executed on the fly, not using static data. We apply this technology to V-doped MoS 2 , uncovering insights into defect formation and evolution under electron beam exposure. This approach opens uncharted avenues for exploring and characterizing materials in dynamic states, offering a pathway to increase our understanding of dynamic phenomena in materials under thermal, chemical, and beam stimuli.

47 OTHER INSTRUMENTATION↗

HopPyBar

HopPyBar is a python program to import, analyze, and export split-Hopkinson pressure bar (SHPB, also known as Kolsky bar) data. Traditional analysis offers a black box approach, where input data is converted to analyzed output by performing a series of calculations without user involvement. This program serves as a developmental platform to "white box" the data analysis process. Data streams can be captured (in-situ) to enable advanced or unconventional analyses, statistics, and comparisons. Additionally, the program is geared towards the standardized forms of input and output used at LANL to streamline analysis, but the open nature of the program makes additional input/output schemes straightforward to add. General workflow will import SHPB data in one of a number of formats, identify relevant portions of data signals, and convert to stress-strain-strain rate to show material behavior as a function of dynamic testing.

Morrow, Benjamin↗

OASIS: Open-Source AI Software Infrastructure for Science-SBIR Phase I

OASIS: Open source AI Software Infrastructure for Science is developed for researchers in the scientific domain. OASIS provides a data API to ingest and serve scientific data formats and annotations within AI workflows. It delivers a unique integration of features such as coupling of data to AI models, scalable training, cloud deployment into a cohesive web and command-line interface, and state-of-the-art techniques to debug and enhance AI models.

Chaudhary, Aashish↗

Conversion Helper 4 Easy Serialization Of Exi (ch4ese)

CH4ESE is an EXI conversion tool developed in Python3 that utilizes the open-source EXIficient implementation of the W3C EXI format specification. CH4ESE can be used to translate to and from EXI format using the command line with input data or using the web server for live-translation.

Rohde, KennethW [Idaho National Laboratory (INL), ↗

A Machine Learning Framework to Deconstruct the Primary Drivers for Electricity Market Price Events

As the electricity grid is moving towards a 100% Renewable Energy Source Bulk Power Grid, the overall operations of the power system operations and electricity markets are changing. The electricity markets are not only dispatching resources economically but also taking into account various controllable actions like renewable curtailment, transmission congestion mitigation, and energy storage optimization to make sure the grid is operating reliably. As a result, price formations in electricity markets have become quite complex. Traditional root cause analysis and statistical approaches are rendered inapplicable to analyze and infer the main drivers behind price formation in the modern grid and markets with variable renewable energy (VRE). In this paper, we propose a machine learning analysis framework to deconstruct some primary drivers for price formation in modern electricity markets with high renewable energy and the outcomes can be utilized for various critical aspects of market design, renewable dispatch and curtailment, operations, and cyber-security applications. The framework can be applied to any ISO or market data and in this paper it is applied to open-source publicly available datasets from California Independent System Operator (CAISO) and ISO New England.

machine learning (ML), electricity markets, Renewa↗