America Chose Its Nuclear Waste Site on a Guess About Cost, not Safety

The U.S. Shouldn’t Make the Same Mistake Twice

America Chose Its Nuclear Waste Site on a Guess About Cost, not Safety

The United States is reconsidering how it sites a nuclear waste repository, and over the past four decades the argument has settled into two camps. One holds that Yucca Mountain was chosen on sound technical grounds and abandoned for political ones; the other holds that the science was always secondary, and that the country should start over where a repository is wanted.

Both camps accept the same assumption: in 1986 the Department of Energy (DOE) ran a real comparison of five candidate sites and got an answer. Unfortunately, DOE’s 1986 decision was flawed. I reconstructed DOE’s 1986 model from every published parameter and replicated the results. Then I tested it in ways that DOE never did. 

Promisingly, DOE’s 1986 evaluations of the five sites were almost perfectly tied on long-term radiological isolation. The rankings between the sites only started to deviate due to both repository cost and transportation cost, accounting for 95.7 percent of all the variation. Those two figures were engineering guesses (that ultimately did not hold up) about building a first-of-a-kind facility in either salt, tuff, or basalt, and were made before anyone had gone underground to evaluate them. Its comparison did not distinguish the sites on the thing a repository must get right—safety—and the numbers proving it were published in the report itself, but never put together. 

By DOE’s own model, all five sites passed the safety test with large margins. What the ranking actually adjudicated, under guidelines that required safety to carry the greater weight, was a pair of unmeasured construction estimates. A decision made on cost is a decision not made on safety.

Deciding on cost among sites that are all safe enough is perfectly defensible. It may even be the right way to decide. But that is not what happened here, because nobody decided it. The analysis produced one combined score per site, and the recommendation was made on those scores. That the score was, in effect, a cost ranking, which was not something DOE set out to do, or that the NRC or Congress could have seen; it does not appear anywhere in the report. It becomes visible only when the model is taken apart and each attribute’s contribution is measured on its own. America chose its nuclear waste site on a cost estimate, by accident.

And the accident landed on the worst possible number. The cost figures were guesses whose published uncertainty was wide enough to reorder the sites in any direction—the least reliable quantity in the model, carrying almost all of the decision. 

If the recommendation had said that cost was choosing the site, the obvious next question would have been whether the cost numbers were firm enough to choose one. They were not.

These are lessons for how the country compares sites, and they must be learned and applied before the next siting choice.

Tied on Safety, Differentiated by Cost

The Nuclear Waste Policy Act of 1982 told DOE to nominate at least five sites and recommend three to the President for characterization—the deep, expensive process of going underground and finding out what was actually there. The December 1984 draft environmental assessments ranked the five against the siting guidelines and combined those ranks with three simple aggregation methods. Many of the 20,000-plus comments that came back went after those methods. In the 1985 review of the draft environmental assessments, the NRC staff suspected one of the failure mechanisms, but did not test it. What they pointed to was preclosure (before sealing) dominating postclosure (long-term isolation) not one metric dominating everything. 

The DOE responded by shifting to a multiattribute utility analysis, published as a separate report in May 1986. That analysis was a standard formal decision model that separated performance measures into the preclosure period and postclosure, and a composite that combined the two. Each of those measures was converted to a common scale and weighted. The weights were elicited, documented, and published, albeit in dense reports.

The analysis measured postclosure, the long-term isolation performance of a geologic repository, at the five sites as nearly identical (Table 1).

Yucca q
Table 1. Postclosure expected utility, reconstructed from DOE/RW-0074, Table 3-6.

The total spread across all five is just 0.23 points out of 100. Four of the five are separated by only 0.01. DOE reported as much itself: unless you assign 99.8 percent of the entire decision to long-term isolation and 0.2 percent to everything else, postclosure does not affect which site comes out on top.

The siting guidelines required that postclosure carry greater weight than preclosure, and in the earlier draft environmental assessments DOE had put it between 51 and 85 percent. But how much postclosure actually weighed was never fixed and the 1986 model set no number at all: it compared every preclosure and postclosure attribute pair one at a time, and published it with DOE’s own caveat that the scaling factors “most emphatically cannot be used as indicators of the importance” of what they weighed.

In one sense, that is reassuring. By DOE’s own 1986 assessment, all five sites would isolate waste acceptably. There is a second reason for the tie: the model scored releases as order of magnitude multiples of the EPA limits. The EPA itself called that figure speculative.

Yucca 2
Figure 1. The scale the 1986 analysis used to score postclosure radiological releases, expressed as a fraction of the EPA release limits over the first 10,000 years. “Significant” is the label the analysis attached to a release at the limit, corresponding, on EPA’s own stated design basis, to no more than one statistical death per decade nationwide. “Statistical death” is a regulatory term that means fatalities estimated to occur among a large population on the basis of the collective dose. 

Every one of the five sites sits at the favorable end of the scale (Figure 1). But it means the ranking—the thing that got used—came entirely from the preclosure model. Reconstructed, it is in Table 2:

Yucca 3
Table 2. Total preclosure utility, reconstructed from DOE/RW-0074, Table 4-21.

The composite ranking is in the same order.

If you put the two tables side by side, the decision’s real structure is in plain view: safety qualified all five sites, and cost picked one. Choosing on cost among sites that all meet the safety standard is a defensible structure, arguably the right one. But nobody said that was the basis. The guidelines required postclosure safety to carry greater weight, and the comparison proceeded as if that weight mattered, when in reality the outcome would have been the same at almost any weight.

Twelve of the Fourteen Measures Did Nothing

Repository cost accounts for 87.9 percent of all the variation separating the five sites’ preclosure ratings. Transportation cost adds another 7.8. The remaining twelve attributes account for 4.3 percent between them.

Yucca 5
Figure 2. Each attribute’s share of the weighted between-site variation, reconstructed from DOE/RW-0074, Tables 4-7 through 4-9.

Think about buying a car. Suppose safety matters more to you than anything else. If every car on the lot carries the same five-star crash rating, that rating cannot help you choose. Safety still matters; an identical number across every option simply cannot separate them, and it makes no difference whether the number is excellent or poor. Only characteristics that actually differ between the cars in front of you matter for your choice: the price, the cargo space, the color. The thing that decides is not the thing you care about most. It is the thing that varies most, weighted by how much you care.

So an attribute can move a ranking only if two things are true at once: the alternatives actually differ on it, and the model places enough weight on it to matter. Fail either test and the attribute is inert, however important it is in principle. The product of the two parameters is the discriminating power, the share of the real difference between the alternatives that this attribute accounts for.

Figure 2 shows what that looks like in practice. On a hundred-point scale, the sites have a spread of 50 points on archaeological and cultural impacts and 19 points on biological impacts. Both are slivers on the chart, because the two attributes together carry a quarter of one percent of the weight. Public non-radiological fatalities are a sliver for the opposite reason: DOE estimated zero public non-radiological fatalities at all five sites, so nothing differentiated them.

Importantly, that means twelve of the fourteen attributes are decision-irrelevant. Take DOE’s own published uncertainty range for each attribute at each site and move any single one of those twelve to the most adverse admissible value for the leading site and the most favorable for its nearest rival, simultaneously, and no ranking changes.

For the two attributes that did matter—repository cost and transportation cost—the DOE had no way to accurately compare the sites. In May 1986, when the analysis was published, not one of the five sites had been fully characterized. This was reflected in the wide ranges the DOE published for cost estimates.

Yucca 6
Figure 3. Reproduced from DOE/RW-0074, Table 4-6. Open circle = DOE’s base-case estimate; bar = DOE’s published range; shaded band = the interval every site’s range contains. Rows in composite-ranking order. These two attributes carry 95.7 percent of discriminating power.

By DOE’s own estimates, any of the five sites could have ended up with the lowest repository cost (Figure 3). A $1.74 billion band of repository cost—$8.38 to $10.12 billion—spans each of the five published ranges. Any site could have cost any amount in that band. The fifth-ranked site coming in cheap and the first-ranked site coming in high would be consistent with DOE’s ranges, not outliers.

Transportation is worse. Every range contains $0.39 to $2.04 billion, a band wider than the entire spread of the base estimates, which run from $0.97 to $1.45 billion.

Yucca 7
Figure 4. The cost error that reverses each pair of sites, computed from DOE’s Table 4-6 estimates. Each percentage is a move in that site’s own base-case cost: the higher-ranked site’s rising, the lower-ranked site’s falling by the same percentage. Dollar swing is the two movements combined, in 1986 dollars.

All ten pairwise rankings reverse within the published cost range (Figure 4). For the closest pair, Richton Dome trades places with Deaf Smith if Richton comes in 3 percent above its estimate while Deaf Smith comes in 3 percent below, well within the ±35 percent error bar. At 6 percent, Yucca Mountain is no longer the top-ranked site: Richton Dome is. Even the widest margin in the model, Yucca Mountain over Hanford, reverses if both are just 23 percent away from the central estimate, still within the cost range.

If every estimate proves optimistic together, the ranking is untouched: DOE’s ranges are roughly proportional to the base estimates, so an error where all the values shift together (i.e., what DOE’s model tested) preserves the ordering and even widens the margins. The reversals come from differential error: one site’s estimate proving optimistic while another’s proves pessimistic. That is the realistic case. These were independent estimates for three different host rocks, produced by different site project offices and their contractors, under assumptions that were never reconciled to a common basis. Even the differences between the estimates, the only thing the model used, sat on inconsistent foundations.

History has since graded the Yucca Mountain cost estimate. DOE’s 1986 base case put building and operating a repository at Yucca Mountain at $7.5 billion in the dollars of the day, the cheapest of the five, inside a published range topping out at $10.1 billion. By 2001 the Department’s life-cycle estimate for the program had reached $57.5 billion; by 2008 it was $96.2 billion in 2007 dollars, $54.8 billion of it for the repository itself. Roughly $15 billion had been spent by the time the program was halted in 2010—no repository built, the license application unresolved. The scopes differ—the later figures span a century and a half of operation and a larger inventory—but no adjustment for inflation closes the gap: the 2008 repository figure is more than three and a half times the inflation-adjusted 1986 base, and well over twice the top of the published range. First-of-a-kind estimates miss, and this one missed the way they usually do: under, by multiples.

Ultimately, the cost estimate was too uncertain to serve as a decision point at all.

DOE Ran a Sensitivity Analysis. It Tested the Wrong Axis.

DOE did not skip sensitivity analysis, and it did not skip the preclosure-postclosure split. The report varies the value tradeoffs within each period, varies the form of the utility function, sweeps the split across its entire range, and re-runs the composite under uniformly optimistic and uniformly pessimistic input. It concludes that the ranking is “insensitive to any reasonable changes in the value judgments or in the form of the utility function.”

That conclusion is correct, but beside the point. The ranking was not sensitive to the value judgments. It was sensitive to the cost estimate that carried almost all of the model’s discriminating power and came with a ±35 percent error bar. And because the analysis moved the inputs only in lockstep—every site was either cheaper or more expensive than expected—the rankings were never going to change. It is in the differential direction—one site proving cheaper, while another more expensive—that the ranking can change, and that direction was never tested.

A one-at-a-time sensitivity analysis on a model whose weight is concentrated in a single uncertain attribute will indicate robustness that the model does not have.

Both the uncertainty ranges and the weights were available in May 1986, eleven pages apart in the same document, but the DOE failed to put them together.

That is less of an indictment of the analysts than it sounds. Varying the weights one at a time was what 1986 computation could afford. Sweeping every input jointly across its published range was out of reach at the time; today it is trivial. The era’s tools could price the value of gathering more information but could not chart where a decision would flip. The analysts tested what could be tested. The underlying problem is that the analysis was not built to ask the question. The DOE robustness test asked whether the ranking survives changes in the evaluators’ values, a question about the decision-makers who provided input. Whether the comparison could resolve the sites at all is a question about the world, and no amount of computing power surfaces a question the architecture never poses. 

Hanford Ranked First or Last

The Secretary recommended Yucca Mountain, Deaf Smith County, and Hanford. In the model those ranked first, third, and fifth. Richton Dome (second) and Davis Canyon (fourth) were passed over.

DOE’s stated reason was diversity of host rock, which is legitimate, as the statute directed sites in different media; three of the five were salt, and Yucca was the only tuff while Hanford was the only basalt. DOE’s recommendation document is candid that the model was not the whole decision: it “does not apply the diversity guidelines required to determine a final order of preference.” As was seen by the postclosure analysis, all five sites met the safety standard with large margins.

But the recommendation offered a second, checkable, defense of Hanford. Hanford, DOE wrote, “scores first in minimizing impacts on the environment, minimizing impacts on socioeconomic conditions.” Luther Carter’s 1987 history records the argument more bluntly: Hanford “actually had ranked tops among the five sites if cost of repository construction and operation and of spent fuel transportation were not taken into account.”

That claim is arithmetically true. Remove both cost attributes and the ranking becomes Hanford, Yucca Mountain, Deaf Smith, Richton Dome, Davis Canyon. Hanford first, exactly as claimed.

The problem is what removing cost removes. Strip the two cost attributes and 95.7 percent of the model’s discriminating power goes with them. The surviving ranking is computed on the remaining 4.3 percent, which were mainly differences the model could not resolve. Hanford’s first place on non-cost attributes and its fifth place overall are not two findings to be weighed against each other. One is a ranking; the other is noise with an ordering imposed on it.

Hanford ranks first whenever cost is given less than 3.7 percent of total weight, and last only once cost exceeds 52.4 percent. Hanford’s position was a function of a single dial, and the analysis reported robustness because it turned a different one.

No site was shown to be better. All five sites met the safety standard, the attributes that mattered could not separate them, and the one that separated them was too uncertain. A comparison this close does not identify a winner; it identifies the need for further information where there is high uncertainty.

Congress Made It Permanent

Eighteen months after the Secretary’s recommendation of three sites to further characterize, Congress amended the Nuclear Waste Policy Act through the Omnibus Budget Reconciliation Act of 1987. The amendments directed characterization only “at the Yucca Mountain site,” ordered site-specific work at every other candidate phased out within 90 days, and terminated the crystalline-rock program that would have produced a second-repository candidate set in the East.

They also added what is now 42 U.S.C. §10172a:

The Secretary may not conduct site-specific activities with respect to a second repository unless Congress has specifically authorized and appropriated funds for such activities.

That provision, requiring both steps, authorize and appropriate, is still law, and it is why no Secretary of Energy—of either party, at any point in thirty-eight years—has been able to study an alternative site.

The 1987 designation was a political act; that much is well understood. But its technical predicate was a comparison in which DOE’s own model scored the sites as tied on safety and separated them by an unmeasured cost estimate.

To be fair to Congress: breaking a tie is what an elected legislature is for, and ex ante Yucca was a defensible pick at the time: it ranked first in the analysis Congress had in hand. Two things turned a legitimate tie-break into the impasse that followed. Nobody told Congress it was a tie, because nobody knew. The comparison arrived as a ranking, under guidelines that said safety carried the greater weight, so Congress could reasonably believe it was ratifying a technical winner rather than making a policy choice. And Congress did more than choose—it abolished the alternatives before the underground characterization that could have confirmed its choice was ever bought. A tie-break picks who goes first; it does not cancel the race.

What a New Comparison Has to Do

The 1986 model was competently built, fully documented, and honestly reported. Its weights were published, its uncertainty ranges were published, its sensitivity analysis was real. Every number in this article comes from DOE’s own tables. What failed was the step between analysis and decision—the moment a ranking with almost no resolving power was handed forward as though it had settled something, and then hardened into statute before anyone checked what it could bear.

That step is not unique to nuclear waste, and it is not behind us. Consent-based siting will require comparing candidate sites, and the comparison will again be built before anyone has gone underground. Three checks would have caught the 1986 failure, and each is cheap enough, and simple enough, to be made a condition of publication for the next comparison.

  • 1. Determine the resolving power of each attribute before reporting a ranking. For every attribute, state its share of the between-site variation, and identify which cannot change the answer anywhere within their own published ranges.
  • 2. Run a sensitivity analysis on the inputs, not just the values. Vary the inputs jointly across their published ranges, in both the common and differential directions. Varying the weights while holding uncertain inputs at base case can overestimate robustness.
  • 3. Report how much error it takes to flip each pair. For each pairwise comparison, state how far the inputs must move to reverse it. In 1986 a 3 percent cost error reversed the closest pair and 23 percent reversed the widest, against a published uncertainty of ±35 percent.

None of this requires new data or a new method. It requires reporting three quantities the analysis already contains.

A deeper, architectural fix is necessary. All fourteen of the 1986 objectives were minimizations—less of everything, forever, with no level at which any of them counted as satisfied. Minimization asks only how low an attribute can go, never whether lower is needed or what it costs, and because the attributes are coupled, pushing one down often pushes others up. A model that assumes lower is always better, no matter how low, cannot say when to stop paying; it converts every attribute into an open-ended purchase. The next comparison should invert that: state thresholds in advance—safety and the lesser impacts as conditions a site either meets or does not—and let cost be the open, declared basis of choice among the sites that meet them. Qualify on the conditions, then choose: that is what the 1986 comparison did in fact, without ever saying so. Saying it in advance would have been honest, and would mean a less expensive overall project, because a threshold defines what is needed and lets the design stop there.

This is not an argument against decision science. It remains the right discipline for siting—the only framework that states out loud what is being traded for what—and the 1986 report is a record of it being used poorly, not of the discipline failing. The tempting alternatives are a probabilistic risk assessment or a full uncertainty propagation, which can be the right tool for licensing a design. But, as a way to choose among sites, it invites a fidelity trap: the ever finer modeling of quantities that cannot change the choice, one more layer of analysis always defensible, the decision waiting on a model that can always be improved. The remedy is not a higher-fidelity model. It is the adjustments above: thresholds stated in advance, resolving power reported before rankings, sensitivity aimed at the inputs that carry the load. A tool is only useful when used well.

Yucca 8
Figure 5. The same five sites under five readings of DOE’s own 1986 numbers; the last column is the ranking that went forward. Rank 1 is the most favorable site, and the shaded blocks are functional ties.

The same defect recurs downstream, in the Yucca Mountain design record, where the decision to include a multi-billion-dollar titanium barrier turned successively on uncertainty reduction, demonstrability, and a classification—none of which can ever return “not needed.”

None of this requires anyone to conclude that Yucca Mountain is unsuitable. The 1986 analysis does not show that, and it does not show the opposite—that is the point. What it shows is that asking the wrong questions gives an answer that isn’t useful.

The next siting study should ask which sites are safe enough, against a stated threshold, and then which of those the country can actually build, with open decisions about cost and consent. Movement on nuclear waste will not begin with a better model. It will begin when the country can again compare more than one place, and knows how to tell whether the comparison means anything.