Evidence synthesis in energy engineering
Energy engineering research asks how much energy a technology can deliver or save and at what cost and emissions. Questions include how much energy retrofits actually save in buildings, how efficiently batteries and hydrogen systems convert and store energy, how heat pumps perform in cold climates, how wind and solar output varies, and how the life cycle emissions of different technologies compare. Reviews are common in this field, and many are narrative or tabulate reported values, though quantitative synthesis with meta-regression is growing.
Engineering evidence has distinctive properties. A reported efficiency is valid only for the stated test conditions. Modelled results differ from measured ones. Field performance depends on installation quality and user behavior. Units and system boundaries vary between studies. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for engineering and technology. The service does not provide engineering design or certification.
Measured versus modelled performance
| Type | Description | Issue for synthesis |
|---|---|---|
| Laboratory test | Controlled test at standard conditions | Upper bound; not representative of installed performance |
| Field measurement | Monitoring of an installed system | Includes installation, climate and user effects; short monitoring periods |
| Simulation | Model prediction from design inputs | Depends on assumptions and calibration; often optimistic |
| Techno-economic model | Cost and performance projection | Depends on prices and learning-rate assumptions |
| Manufacturer data | Reported specifications | Test conditions chosen by the supplier; may differ from use |
Reviews of building energy efficiency have documented a performance gap between predicted and measured energy use, with measured savings from retrofits often smaller than predicted. The gap has been attributed to modelling assumptions, workmanship, and occupant behavior, including the rebound effect, in which more efficient services are used more. A review should keep the types of values separate, code the test conditions, and report the ratio of measured to predicted values where both exist. Pooling measured and modelled values in one estimate without a type moderator would mix different quantities.
Test conditions, units and system boundaries
An efficiency number is meaningful only with its conditions: temperature, load, state of charge, cycling protocol and the system boundary (cell, module or whole system, including auxiliary loads). Round-trip efficiency of a battery measured at the cell level differs from that of the installed system. Solar module efficiency under standard test conditions differs from annual yield. A review should extract these conditions, convert to common units, and compare only like with like. When important information is missing, the review says so and either contacts authors or excludes the value, with a recorded reason.
Technology generations matter. Efficiencies and costs change quickly, and a value from ten years ago may not describe current products. Publication year or technology readiness level should be coded as a moderator.
Buildings, behavior and rebound
Studies of energy-saving programs for households and buildings include randomized trials of feedback and behavioral nudges, quasi-experiments on subsidies and standards, and engineering evaluations of retrofits. Average savings from home energy reports are small, a few percent, but consistent across large trials. Retrofit savings vary with climate, building type and baseline use. Rebound effects reduce the savings from efficiency measures, and estimates of their size vary with the method and the service studied. A review should code the program type, baseline consumption, measurement approach and duration, and should separate evidence from large randomized trials from smaller engineering studies. The environmental economics page describes related methods for elasticities.
Storage, conversion and degradation
Studies of batteries, hydrogen systems and thermal storage report capacity, efficiency and degradation. Degradation is measured by cycling in the laboratory under accelerated conditions that may not match real use, and field data on long-term degradation are scarcer. Reviews should record cycle depth, rate, temperature and the definition of end of life, and avoid projecting lifetimes beyond the data. Chemistry and design vary greatly, and a pooled value across chemistries is rarely meaningful. A systematic map of tested conditions often serves better than a pooled average.
Life cycle assessment and harmonization
Life cycle assessments estimate emissions or other impacts over a product's life, and published values for the same technology vary widely because of different system boundaries, assumptions about electricity mix, lifetimes and allocation methods. Harmonization studies adjust reported values to common assumptions, narrowing the spread, and then summarize the result. A review should record the assumptions of each study, apply harmonization transparently and show the distribution before and after. Results from one region's electricity mix do not describe another's. Reviews of life cycle studies are described further under environment and sustainability.
Generation, variability and system studies
Studies of wind, solar and grid integration use measured production data, weather models and system simulations. Capacity factors vary by site and year, and projections for new sites depend on models. System-level studies of how much variable generation a grid can integrate depend on assumptions about flexibility, storage and demand, and they are model outputs with scenario dependence. A review should describe scenario assumptions, avoid pooling scenarios as if they were observations and state the sensitivity of conclusions to key assumptions.
Publication bias and optimism
New technologies are reported by their developers, in best-case conditions, and failed pilots are seldom published. We compare laboratory with field results, independent with developer-led tests, and early with later studies of the same technology. Optimism in cost projections has been documented, and a review of cost studies should compare past projections with realized costs where possible.
Costs, learning curves and projections
Cost data for energy technologies are reported as cost per kilowatt of capacity, per kilowatt-hour delivered or as levelized cost of energy. The levelized cost depends on the discount rate, lifetime, capacity factor and fuel price assumptions, which differ between studies and are often not stated in full. Learning curves relate cost to cumulative production and are used to project future costs. Past projections for several technologies have been both too pessimistic and too optimistic, depending on the technology and the period. A review of cost evidence should extract the assumptions behind each value, convert currencies and years as described on the health economics page, and show how much of the spread in reported costs is explained by the assumptions.
Subsidies, taxes and local conditions affect observed costs, and studies differ in whether they include them. We record this and avoid comparing a subsidized price with an unsubsidized one, and we state the year and the market for every price reported.
Safety, reliability and failure data
Reliability studies report failure rates, mean time between failures and degradation of components in service. Failure data are often proprietary or come from small samples, and survivorship bias is a risk if only systems still in operation are monitored. Definitions of failure differ, and data from accelerated tests do not map simply onto field lifetimes. A review should state these limits, treat failure rates as counts with exposure time and use appropriate rate models, as described on the incidence meta-analysis page. Safety assessment of any installation is a task for qualified engineers and regulators and is outside this service. We also record whether failure data came from the manufacturer, an operator or an independent test body, since the source affects how complete and unbiased the records are likely to be, and we report the exposure time and number of units behind every rate we cite.
Common pitfalls we look for
- Pooling laboratory, field and modelled values.
- Comparing efficiencies at different test conditions or system boundaries.
- Ignoring technology generation and year.
- Treating predicted savings as realized savings.
- Treating scenario outputs as observations.
- Projecting lifetimes beyond tested conditions.
Planning an energy engineering synthesis
We help define the technology, setting and performance measure, plan searches in Scopus, Web of Science, IEEE Xplore, Engineering Village and agency reports, and set up coding of test conditions, system boundary, location, climate, year, measurement duration and data type. See the systematic review service for scope and process.
An invented example of a performance gap
Suppose a review of 40 invented building retrofits finds an average predicted saving of 30 percent in heating energy and an average measured saving of 21 percent. The ratio of measured to predicted is 0.70. If a prediction interval for the ratio runs from 0.40 to 1.00, then even though the average gap is large, some buildings met the prediction. A planner using the predicted figure to estimate savings would overestimate on average by about 9 percentage points, or by about 43 percent of the realized saving (9 divided by 21). The review would report the ratio and its range and explain what factors were associated with larger gaps.
Coding and transparency
Coding frames record technology and design, test or monitoring conditions, system boundary, units, location and climate, year and technology readiness, data type (measured, modelled, manufacturer), duration, sample size (units monitored) and funding. Values from figures are digitized and checked by a second coder. Two coders work independently on a sample and agreement is reported, and the coded data and scripts are shared with the final report.
How we support research projects in this area
From a technical question to a published review
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
Help choosing between a mapping study, a systematic review and a meta-analysis, and a protocol with research questions.
Searching and extraction
Database search, snowballing and grey literature, with screening and a classification scheme.
Synthesis
Structured synthesis, quality assessment and, where experiments are comparable, meta-analysis.
Manuscript and submission
The manuscript, the replication package and journal preparation.
Boundaries of this service
An energy engineering synthesis describes published measurements and models. It does not provide engineering design, system sizing, safety certification or investment advice. Performance depends on conditions, installation and use, and technology changes quickly, so published values may not describe current products or a particular site.
Frequently asked questions
Why do measured savings differ from predicted ones?
Because of modelling assumptions, workmanship, occupant behavior and rebound effects. Reviews report the ratio of measured to predicted values.
Why can't efficiencies be compared directly?
Efficiency depends on test conditions and system boundary, so only values measured under comparable conditions are compared.
What is harmonization in life cycle studies?
Adjusting reported results to common assumptions so that differences in methods do not hide real differences between technologies.
Can scenario outputs be meta-analyzed?
No. They are model results that depend on assumptions, so they are described and compared, not pooled as observations.
How is technology change handled?
Year and technology readiness are coded, and findings are dated so readers know which generation they describe.
Do you provide engineering design or certification?
No. The service provides research and evidence-synthesis support only.
References
- van Dronkelaar C, Dowson M, Burman E, Spataru C, Mumovic D. A review of the energy performance gap and its underlying causes in non-domestic buildings. Front Mech Eng. 2016;1:17.
- Allcott H. Social norms and energy conservation. J Public Econ. 2011;95(9-10):1082-1095.
- Sorrell S, Dimitropoulos J, Sommerville M. Empirical estimates of the direct rebound effect: a review. Energy Policy. 2009;37(4):1356-1371.
- Warner EJ, Heath GA. Life cycle greenhouse gas emissions of nuclear electricity generation: systematic review and harmonization. J Ind Ecol. 2012;16(S1):S73-S92.
- Sovacool BK, Axsen J, Sorrell S. Promoting novelty, rigor, and style in energy social science: towards codes of practice for appropriate methods and research design. Energy Res Soc Sci. 2018;45:12-42.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.