Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “PV data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Relating Aerial Infrared Thermography Defects to Photovoltaic Performance: Preprint

In this research, we examine the relationship between aerial IR defect analysis and photovoltaic (PV) performance data for twelve utility- and commercial-scale solar sites in the United States. To do this, we fuse the site diagram geoJSON's, aerial infrared thermography (aIRT) defect analyses, and associated inverter time series, allowing for a direct comparison between site defects and time series data. Defect analyses were provided by Zeitview, under its Solar Insights platform. Following the data fusion process, we look at the relationship between system performance and aIRT defects. We investigate the relationship between degradation and hotspot defects, as well as the relationship between AC power data and offline strings and misaligned modules. In general, system degradation was not affected by long-term or balance-of-system (BoS) defects as they occurred infrequently in the data set. However, for one system, a near statistically significant relationship (p-value=0.057) was found when comparing the degradation of inverter blocks with several multi-hotspot defects to all other inverter blocks without this particular defect. There was strong alignment when comparing short-term recoverable module defects such as stuck trackers and offline strings to time series data. In general, we found that when an inverter block has more than 80% of modules flagged for one of these defects, its AC power time data is flat-lined and the inverter block is not producing.

aerial inspection↗

Integrated Large-Scale Data Management Platform for Photovoltaic Power Conversion Equipment (PCE) Reliability Data

To meet the demand for accuracy and real-time capability of PV system degradation evaluation, massive volume data is needed to run high-fidelity and high-efficiency simulations and perform advanced data analysis. However, PV farm operators have a series of difficulties with PV inverter data, such as data collection from multiple channels, massive data storage, data management and massive data analysis. To address these challenges, we developed an integrated data management platform capable of data acquisition, processing, storage, query, and performing big data analysis utilizing AI algorithms. The platform can also achieve data correctness verification and provide an effective distributed data management solution to retrieve massive data and establish a connection to distributed computational frameworks.

data management platform↗

Integrated Large-Scale Data Management Platform for Photovoltaic Power Conversion Equipment (PCE) Reliability Data: Preprint

To meet the demand for accuracy and real-time capability of PV system degradation evaluation, massive volume data is needed to run high-fidelity and high-efficiency simulations and perform advanced data analysis. However, PV farm operators have a series of difficulties with PV inverter data, such as data collection from multiple channels, massive data storage, data management and massive data analysis. To address these challenges, we developed an integrated data management platform capable of data acquisition, processing, storage, query, and performing big data analysis utilizing AI algorithms. The platform can also achieve data correctness verification and provide an effective distributed data management solution to retrieve massive data and establish a connection to distributed computational frameworks.

data management↗

Phase Identification in Real Distribution Networks with High PV Penetration Using Advanced Metering Infrastructure Data

Many distribution network monitoring and control applications - including state estimation, volt/VAR optimization, and network reconfiguration - rely on accurate network models; however, the network models maintained by utilities can become outdated because of restoration activities, network reconfiguration, and missing data. With the widespread deployment of advanced metering infrastructure (AMI), abundant measurement data from low-voltage secondary networks are available. The AMI measurement data can be used for phase identification to improve the network models. Although the existing phase identification techniques work well in passive distribution feeders that do not have photovoltaic (PV) generation, they can fail to accurately identify the phases in the presence of PV. This paper proposes a robust phase identification algorithm based on supervised machine learning that accurately identifies the AMI meter phase connectivity in the presence of significant PV generation. The proposed algorithm does not require network topology information or feeder head measurement data. The algorithm is validated using the AMI measurement data collected in the field and the field-validated phase connectivity database on two real distribution feeders from San Diego Gas & Electric Company that have significant PV generation.

advanced metering infrastructure↗

Phase Identification in Real Distribution Networks with High PV Penetration Using Advanced Metering Infrastructure Data: Preprint

Many distribution network monitoring and control applications - including state estimation, volt/VAR optimization, and network reconfiguration - rely on accurate network models; however, the network models maintained by utilities can become outdated because of restoration activities, network reconfiguration, and missing data. With the widespread deployment of advanced metering infrastructure (AMI), abundant measurement data from low-voltage secondary networks are available. The AMI measurement data can be used for phase identification to improve the network models. Although the existing phase identification techniques work well in passive distribution feeders that do not have photovoltaic (PV) generation, they can fail to accurately identify the phases in the presence of PV. This paper proposes a robust phase identification algorithm based on supervised machine learning that accurately identifies the AMI meter phase connectivity in the presence of significant PV generation. The proposed algorithm does not require network topology information or feeder head measurement data. The algorithm is validated using the AMI measurement data collected in the field and the field-validated phase connectivity database on two real distribution feeders from San Diego Gas & Electric Company that have significant PV generation.

advanced metering infrastructure↗

Phase Identification in Real Distribution Networks with High PV Penetration Using Advanced Metering Infrastructure Data

Many distribution network monitoring and control applications - including state estimation, Volt/VAr optimization, and network reconfiguration - rely on accurate network models; however, the network models maintained by utilities can become outdated because of restoration activities, network reconfiguration, and missing data. With the widespread deployment of advanced metering infrastructure (AMI), abundant measurement data from low-voltage secondary networks are available. The AMI measurement data can be used for phase identification to improve the network models. Although the existing phase identification techniques work well in passive distribution feeders that do not have photovoltaic (PV) generation, they can fail to accurately identify the phases in the presence of PV. This paper proposes a robust phase identification algorithm based on supervised machine learning that accurately identifies the AMI meter phase connectivity in the presence of significant PV generation. The proposed algorithm does not require network topology information or feeder-head measurement data. The algorithm is validated using the AMI measurement data collected in the field and the field-validated phase connectivity database on two real distribution feeders from San Diego Gas & Electric Company that have significant PV generation.

advanced metering infrastructure (AMI)↗

Transient Data Library of Solar Grid Integrated Distributed System

This submission contains an open-source library of transient events in distributed system with high solar PV. The library includes the collected data, related documents and scripts for loading the data. The data library is built for transient event detection and machine learning based analysis algorithm development. The data was collected via both field test and software simulation. The units for the data are included in the data file headers for each data series. A text editor or spreadsheet software, such as Excel, and Matlab is required to view the data.

algorithms↗

A novel data gaps filling method for solar PV output forecasting

This study proposes a modified gaps filling method, expanding the column mean imputation method and evaluated using randomly generated missing values comprising 5%, 10%, 15%, and 20% of the original data on power output. The XGBoost algorithm was implemented as a forecasting model using the original and processed datasets and two sources of solar radiation data, namely, Shortwave Radiation (SWR) from Advanced Himawari Imager 8 (AHI-8) and Surface Solar Radiation Downward (SSRD) from ERA5 global reanalysis data. Further, the accuracy of the two sets of forecasted power output was evaluated using Root Mean Square Error (RMSE) and Mean Absolute Error (MAE). Results show that by applying the proposed gap filling method and using SWR in forecasting solar photovoltaic (PV) output, the improvement in the RMSE and MAE values range from 12.52% to 24.30% and from 21.10% to 31.31%, respectively. Meanwhile, using SSRD, the improvement in the RMSE values range from 14.01% to 28.54% and MAE values from 22.39% to 35.53%. To further evaluate the accuracy of the proposed gap-filling method, the proposed method could be validated using different datasets and other forecasting methods. Future studies could also consider applying the said method to datasets with data gaps higher than 20%.

Energy & Fuels↗

Understanding Solar Photovoltaic System Performance: An Assessment of 75 Federal Photovoltaic Systems

This report presents a performance analysis of 75 photovoltaic systems based on PV system production data collected as part of a FEMP Federal PV Performance Assessment project combined with co-incident insolation, and ambient temperature to analyze how actual performance compares with a performance model. FEMP collaborated with 17 Federal agencies and sub-agencies to collect the information required to analyze the performance of each system. The systems represent a total capacity of 30,714 kW and range in size from 1 kW to 4,043 kW, with an average size of 410 kW, and were installed between 2011 and 2020. The data is analyzed for Key Performance Indicators, Availability, Performance Ratio and Energy Ratio by comparing the measured production data to model production data. The System Advisor Model (SAM) combines a description of the system (such as inverter capacity, de-rating for temperature, balance-of-system efficiency) with environmental parameters (coincident solar and temperature data) to calculate predicted performance. The performance metrics are calculated by lining up the measured production data with the model estimate on an hour-by-hour, day-by-day, or month-by-month basis (depending on the interval resolution of the production data). A report with system description, photo of the system, special assumptions made for the site, graph of measured production and model production, table of key performance indicators, and links to O&M resources that might improve performance was produced and delivered to site and agency staff with a short on-line briefing.

14 SOLAR ENERGY↗

PV_ICE: PV in the Circular Economy, Dynamic Energy and Materials Tool (PV ICE)

This open-source tool implements Circular Economy metrics for photovoltaic (PV) materials. It can be used to quantify and assign a value framework to efforts on re-design, reduction, replacement, reuse, recycling, and lifetime and reliability increases in the PV value chain. The PV_ICE is leveraging published data from different sources on PV manufacturing and predicted technological changes. Input data is being compiled here This tool will help implement circularity metrics, quantify and assign a value framework to efforts on re-design, reduction, replacement, reusage, recycling, and lifetime and reliability increases on PV.

Ovaitt, nee Ayala Pelaez, Silvana↗

Optimal PV Inverter Control in Distribution Systems via Data-Driven Distributionally Robust Optimization

Distribution systems with high penetration of uncertain solar generation call for advanced control strategies of photovoltaics (PVs) inverters. This paper proposes a data-driven distributionally robust optimization (DDDRO) approach to optimally controlling the PV inverters to improve the system operation performance under solar power uncertainties. In the proposed DDDRO approach, a Wasserstein ball-based method is proposed to construct the distributional ambiguity set to model the uncertainties of PV generation through partial observations of historical data without knowing exact probability distributions. We further reformulate the computationally intractable DDDRO model to a mixed integer second order cone programming (MISOCP) problem. The effectiveness and out-of-sample performance of the proposed approach have been demonstrated on a modified IEEE 33-node system. We conduct a comparative study to compare the proposed method with traditional chance constrained programming (CCP). It shows that the proposed DDDRO approach can provide a less conservative yet robust solution to minimize the worse-case expectation of the total network loss while maintaining nodal voltages in a secure range.

Xue, Yaosuo↗

PV Validation Hub

The Validation Hub will be a clearinghouse for the transfer of novel algorithms and software from the PV research community to industry. Potential algorithms tested in the Hub could include the estimation of various PV loss factors and the detection of various operational issues. The primary function of the Hub will be for developers to submit executable code which will run on hosted data sets. Developers will receive private reports on the accuracy and performance (e.g., run-time) of the submitted algorithms, and public high level summaries will be hosted. These summaries will indicate the organization who submitted the algorithm (e.g., links to GitHub pages, documentation websites, etc.), high-level accuracy metrics, and standardized performance metrics. These results will be stored in a publicly available database, accessible through the Hub, with the ability for users to sort and filter the results. In short, the Hub will be presented to public users as a collection of interactive leaderboards, organized around specific analysis tasks pertinent to the PV data science community. These tasks include things such as the estimation of various PV loss factors and the detection of various operational issues. We will present progress on the development of this hub, including preliminary results of comparative validation of PV data science algorithms and progress towards building the platform itself.

algorithm↗

Extending Component Lifetime And Improving Inverter Reliability (ECLAIIR)

Inverter reliability remains one of the most persistent challenges limiting the performance, availability, and economic viability of utility‑scale photovoltaic (PV) plants. Industry data consistently show that inverters account for the highest share of corrective maintenance events and unplanned outages across PV fleets. These failures result in energy losses, increased O&M costs, and reduced confidence in long‑term solar asset performance. Motivated by these challenges, this project—Extending Component Lifetime and Improving Inverter Reliability (ECLAIIR)—was undertaken to systematically investigate inverter degradation and failure mechanisms, develop predictive maintenance capabilities, and establish data‑driven pathways to improve service life and reduce the Levelized Cost of Energy (LCOE) for large‑scale PV systems. The primary goal of the project was to identify pre‑failure signatures in string inverters using both lab‑based accelerated lifetime testing and field‑based data and to develop predictive maintenance algorithms that can anticipate inverter faults before they occur. Through collaboration with inverter testing laboratory, solar PV plant owner, and failure‑analysis experts, the project advanced the technical understanding of inverter reliability. By instrumenting inverters with thermistors, humidity sensors, power‑quality meters, and acoustic sensors, the research established how multiple sensing modalities can reliably detect deviations from normal behavior hours to days before failure. These findings substantially enhance scientific understanding of inverter failure kinetics and provide the PV industry with the most comprehensive cross‑OEM characterization of early‑stage failure indicators reported to date. Technically, the project demonstrated the effectiveness of predictive maintenance by developing and validating the PreDICT (Predictive Diagnostics of PV Inverters Using Condition Monitoring and Trend Analysis) framework—a multi‑layer diagnostic architecture combining peer‑to‑peer analytics, historical trend modeling, and advanced machine‑learning techniques such as the Sequential Conditional Variational Autoencoder (SCVAE). This predictive model achieved more than 90% accuracy in detecting pre‑failure conditions and provided up to four days of lead time before inverter failure in field scenarios. Economically, the project’s LCOE analysis showed that predictive maintenance can reduce lifetime energy losses and minimize corrective maintenance interventions. Modeling indicated that, depending on inverter failure rates and replacement timelines, predictive maintenance can significantly reduce LCOE impacts associated with inverter downtime: from as high as 19.4% under conventional maintenance strategies to 0.1%–10.17% when predictive analytics are adopted. These results confirm that predictive maintenance is both technically feasible and economically advantageous for utilities and plant operators. The project’s findings also have broad public benefit. By improving inverter reliability and reducing downtime, predictive maintenance directly increases electricity generation from existing PV assets. Enhanced reliability lowers operational costs for utilities, which can translate over time into lower energy costs for consumers. Furthermore, the project’s technical publications, conference presentations, and industry workshops ensure that knowledge gained is shared broadly across the solar industry, supporting workforce development and enabling utilities of all sizes to adopt modern asset‑health monitoring practices. The retrofitting case study and service‑life prediction framework further support informed decision‑making for aging PV fleets, helping operators extend system life and reduce electronic waste. In summary, the ECLAIIR project significantly advanced the state of knowledge on inverter degradation, demonstrated the technical and economic value of predictive maintenance, and delivered actionable tools and insights that support more reliable, cost‑effective, and sustainable PV plant operation. The outcomes of this project will continue to inform utility practices, guide inverter design improvements, and strengthen the long‑term performance of solar assets nationwide.

14 SOLAR ENERGY↗

PVplr-stGNN 0.1.10

PV Performance Loss Rate Estimation using Spatio-temporal Graph Neural Networks PVplr-stGNN is a Python 3 package developed by the SDLE Research Center at Case Western Reserve University in Cleveland OH. This repository contains the full source PVplr-stGNN package. The package contains the PV-stGAE for missingness data detection and imputation and PV-DynGNN for PLR estimation.

Fan, Yangxin [Case Western Reserve Univ., Clevelan↗

Public Reference Data for Megawatt-Scale Hydrogen Electrolysis - NLR Historical Solar PV

The U.S. Department of Energy and National Laboratory of the Rockies (NLR) demonstrate hydrogen electrolysis from variable sources, hydrogen compression and storage, and hydrogen fuel cell power production using megawatt-scale equipment at NLR’s Flatirons Campus as part of the Advanced Research on Integrated Energy Systems (ARIES) research platform. This dataset represents part of that effort and is intended for academic, national laboratory, industrial, and other stakeholders to plan, design, and validate models of megawatt-scale hydrogen technologies and diverse energy infrastructure nationwide. These data provide a baseline for how existing hydrogen electrolysis technologies perform when coupled with various energy technologies. Future datasets will demonstrate how existing hydrogen fuel cell technologies can provide controllable, dispatchable, and variable power output for artificial intelligence data centers and other variable loads. This dataset entry describes the behavior of a 1.25-MW proton exchange membrane MC250 electrolyzer system, manufactured by Nel Hydrogen , [1] when fed historical data generated by the 430-kW, fixed-axis solar photovoltaic (PV) array located at NLR’s Flatirons Campus. (While the electrolyzer balance of plant supports up to 2.5 MW of electrolysis, NLR only has a single 1.25-MW electrolysis stack.) Solar PV power output data for the 2020 calendar year were categorized on a daily basis by total energy generation and standard deviation. Each day was then ranked by these metrics, and the 25th, 50th, and 100th percentiles were selected. The 75th percentile day did not exhibit sufficient variability to make for a valuable experiment. A similar process was used for the related historical wind dataset . [2] The historical days in 2020 that represented these percentiles are Dec. 19, March 29, and May 4, respectively. The entire solar day’s power profile was then fed through the MC250 electrolyzer. Due to its length, the 100th percentile day experiment was split into two parts, and the final 3 hours of the solar day were not captured. These final 3 hours contained no spikes or dips of interest and simply represented a slow decay of input solar power. Also, a single timestamp (13:13:47 on Jan. 14, 2026) was lost in the hydrogen system supervisory control and data acquisition. Finally, during the 25th percentile experiment (solar day Dec. 19, 2020) data recording was lost from 11:00:13 to 11:14:45. The roughly 15 minutes of the solar profile were rerun at the end of the experiment and spliced into this time slot during post-processing. The electrolysis system controls hydrogen production by varying direct current applied to the stack, from a maximum of 3,000 A to a minimum safe operation of 300 A, or 10%. Because the current–voltage characteristic changes as the stack ages and efficiency degrades, the actual minimum safe operating power changes over time. The historical solar profiles were translated from power (kilowatts) to current (amperes) using a curve fit with calibration data and sent to the electrolyzer power supply at 1-Hz frequency. For more details on the statistical analysis process, see the slide deck “Public Reference Data for Megawatt-Scale Hydrogen Electrolysis: NLR Historical Solar PV Analysis and Profile Generation” accessible with this data entry. These datasets report relevant hydrogen balance-of-plant and system data, all captured at 1 Hz, including hydrogen mass production measured with an Emerson Coriolis flow meter. Each .zip file represents a single solar PV electrolysis experiment and is formatted as: {technology}_{percentile}_{scaling factor} For instance, “solarPV-430kW_25_2x.zip” reports the experiment using the 25th percentile solar data from the historical 2020 solar PV dataset, scaled to 200%. Scaling factors were applied to the generated solar PV power output files to more closely match the 1.25-MW capacity of the electrolyzer. Each .zip folder contains the following files: A .csv file containing raw data. An .xlsx file explaining all the fields in the raw data. A .png plot showing the time series of hydrogen production, electrolysis power consumption, and solar power input. A PDF file detailing the historical solar data statistical analysis used to generate the solar profile. An experiment labeled “characterization_200.zip” demonstrates the MC250 electrolyzer steady-state response with 30-minute load steps for a total duration of 5 hours. Finally, a .csv file is provided with all experiments combined into one dataset labeled "combined_solarPV_experiments.csv". [1] nelhydrogen.com/product/mc-series-electrolyser . [2] data.nlr.gov/submissions/316 .

08 HYDROGEN↗

Satellite Imagery of PV Site Storm Damage

"This repository contains multiple data sets focused on visible damage to photovoltaic (PV) installations following extreme weather events such as hailstorms and hurricanes. Data sets are split into two categories: the first category, the ‘manually labeled’ data, was compiled by researchers manually, and contains manually identified PV sites exposed to storms. The second data set, the ‘aggregated’ data, is a compilation of the manually labeled PV sites and deep learning-identified PV sites. The hail damage data set focuses on post-storm PV damage following a September 24, 2023 hailstorm in Austin, TX, which caused over $600 million in damages in the Austin metro area. The hurricane damage data set focuses on post-storm PV damage following Hurricanes Irma and Maria in Puerto Rico and the US Virgin Islands. Hurricanes Irma and Maria were back-to-back category 5 hurricanes, which pummeled the Caribbean and southeastern United States in September 2017, causing an estimated $115.2 billion in damages."

14 SOLAR ENERGY↗

Numerical Validation of an Algorithm for Combined Soiling and Degradation Analysis of Photovoltaic Systems

We describe and demonstrate an open-source algorithm for simultaneously quantifying degradation and soiling of photovoltaic (PV) systems from energy-production time series data. The new analysis is based on year-on-year degradation rate analysis combined with stochastic rate and recovery soiling analysis. The algorithm is designed to fit into the workflow provided by RdTools, a Python module maintained by NREL and collaboratively developed with the community, which provides a framework and functions for degradation and loss-factor analysis of PV field data. We demonstrate the method on numerically simulated PV data sets and show that it reduces the root-mean-square error of the P50 degradation rate estimate when soiling is present.

14 SOLAR ENERGY↗

How Climate and Data Quality Impact Photovoltaic Performance Loss Rate Estimations

Different data pipelines and statistical methods are applied to photovoltaic (PV) performance datasets to quantify the performance loss rate (PLR). Since the real values of PLR are unknown, a variety of unvalidated values are reported. As such, the PV industry commonly assumes PLR based on statistically extracted ranges from the literature. However, the accuracy and uncertainty of PLR depend on several parameters including seasonality, local climatic conditions, and the response of a particular PV technology. In addition, the specific data pipeline and statistical method used affect the accuracy and uncertainty. To provide insights, a framework of (≈200 million) synthetic simulations of PV performance datasets using data from different climates is developed. Time series with known PLR and data quality are synthesized, and large parametric studies are conducted to examine the accuracy and uncertainty of different statistical approaches over the contiguous US, with an emphasis on the publicly available and “standardized” library, RdTools . In the results, it is confirmed that PLRs from RdTools are unbiased on average, but the accuracy and uncertainty of individual PLR estimates vary with climate zone, data quality, PV technology, and choice of analysis workflow. Best practices and improvement recommendations based on the findings of this study are provided.

14 SOLAR ENERGY↗