FAIRification, Quality Assessment, and Missingness Pattern Discovery for Spatiotemporal Photovoltaic Data
The ongoing growth of the photovoltaic market has pushed the demand for power forecasting and performance evaluation for a huge population of PV power plants. Through access to a large number of time series data sets from different power plants, we have found common issues that impede the modeling process. Namely, the time series data are hard to transfer between groups due to differences in variable nomenclature, and the quality of the data sets can vary. We address the issue of variable nomenclature by FAIRifying spatiotemporal PV time series data. Through the creation of a solar power plant ontology, we propose standards for the naming and structure of metadata used to describe the data from these power plants. Using the structure from this ontology, we have developed both R and Python packages for the automation of the FAIRification process. We have also developed an R package that automates the analysis of the quality of a data set through the designation of letter grades. With access to large time series data sets across many power plants, we can utilize spatiotemporal coherence between the sites in order to improve the quality of our data. To solve the issue of data missingness, we propose the use of Spatiotemporal-GNN autoencoders to detect and impute missing values from a data set by utilizing data from power plants nearby.