Article Open Access

Dynamic Differential Games of Nuclear Proliferation and Preemption under Uncertainty

Innocent Adinebo Katule Innocent Adinebo Katule Google Scholar More articles 1 Akindele Michael Okedoye Akindele Michael Okedoye Google Scholar More articles 1

  1. 1Department of Mathematics, College of Science, Federal University of Petroleum Resources, Effurun, Nigeria; 330102
Dialogic STEM Journal · Vol 1, Issue 1 · 5 September 2026 ·pp. 1-32 · Open Access (CC BY 4.0) · 250 66

https://doi.org/10.66845/dstem.2026.00003

Download PDF
Share
Contents

Abstract

The nuclear proliferation problem is an important issue in the field of nuclear strategy, characterized by dynamic in-vestments by states and possibilities of preventive actions. The conventional static game models cannot capture dynamic processes, uncertainties, and irreversibility inherent in nuclear strategy problems. This work aims to develop a unified differential game model for nuclear strategy problems, addressing the gaps identified by conventional models. Specif-ically, we propose a comprehensive differential game model for nuclear strategy problems, focusing on the interaction between a proliferating state, e.g., Iran, and a preventive state, e.g., Israel/US, within the context of nuclear weapons development. Our differential game model combines four distinct yet integrated models, each of which progressively incorporates all critical elements of nuclear strategy problems. Model 1: This model establishes a one-state differential game foundation, where the proliferator invests resources in nuclear capability, and the preventer invests resources in prevention, including a costly, capability-reducing preventive strike. Model 2: This model incorporates asymmetric in-formation, where the speed of nuclear breakout by the proliferator is considered a private information, and the preventer learns from an imperfect observation of the state through Bayesian updating. Model 3: This model employs a multi-stage, continuous time Markov chain method, capturing dynamic nuclear R&D, enrichment, and weaponization, as well as the stage-dependent effectiveness of prevention. Model 4: This model develops a two-state differential game model, where the proliferator has the option of investing resources in deterrence capabilities, making a preventive strike by the pre-venter more costly. For each model, we derive the corresponding Hamilton-Jacobi-Bellman (HJB) equations, and then analyse the optimal feedback strategies, as well as the equilibrium conditions, for each model. Our results reveal im-portant strategic phenomena, including signalling incentives, pre-emption traps, and deterrence stability, providing a rigorous mathematical foundation for nuclear strategy problems, with testable implications for policy, e.g., optimal sanctions, intelligence, and military intervention.

Keywords

Differential games Nuclear proliferation Deterrence Game theory Uncertainty

Full text

Reading view
PDF version

1. Introduction

The decision of a state to seek nuclear weapons and the parallel decision of an adversary to engage in preventive action to dissuade or deter the acquisition of such weapons is one of the most important strategic interactions in international politics (Sagan & Waltz, 2013). A core aspect of the problem is the proliferation prevention dilemma, in which the proliferator seeks the ultimate form of deterrence, while the preventer is worried about the impact on regional or global power balances. Several historical examples, such as the 1981 Osirak reactor strike and the more contemporary Iran negotiations, highlight the complexity of the problem, which includes substantial investment incentives, escalating tensions, and the ongoing threat of military action (Feldman, 2011; Parsi, 2017).

Theoreticians have also contributed valuable insights to the problem. Initial attempts to address the problem have utilized static or two-player game theoretic approaches to analyse threat credibility and conditions for proliferation (Bueno de Mesquita & Riker, 1982; Powell, 1990). However, such approaches have the disadvantage of failing to account for the temporal dimension of the problem, which is crucial in determining the evolution of the situation. More contemporary approaches have utilized dynamic game theoretic approaches. Bas & Coe (2016) show through the application of a dynamic signalling game model the potential for a state to slow down proliferation to enhance the credibility of deterrence. Narang (2022) also examines the “proliferation rings” and their dynamic implications.

Operational research and game theory provide the relevant tools for this dynamic analysis. The use of differential games in the context of an arms race was first explored in Simaan & Cruz (1975), followed by Brito & Intriligator (1985). More recently, Acemoglu & Wolitzky (2014) employed the dynamic game approach to examine the economics of conflict and appropriation. Yet, these models often model conflict as a continuous process, not a discrete event, which might or might not occur, like a preventive strike. The real options approach, which was first employed by Dixit & Pindyck (1994), has been used to examine investment decisions under uncertainty. The approach was recently employed to examine the decision to initiate war, which was modelled as an optimal stopping problem (Baliga & Sjöström, 2004; Meirowitz & Sartori, 2008). Despite these advances, there is still a lack of a unified framework that simultaneously accounts for the major elements of the proliferation problem, the dynamics of endogenous growth, costly prevention, the optionality of preventive strikes, asymmetric information, multi-stage technological hurdles, and deterrence. The current paper proposes a comprehensive framework for the nuclear proliferation problem, which is composed of a series of four-part differential games.

The methodology employed is gradual, meaning that the analysis begins with a simple one-dimensional model that identifies the trade-offs between investment and prevention. Complexity is then added to the model, which mirrors the real-world context. In Model 1, the basic framework is established, which includes continuous state and control variables, as well as an optimal stopping problem for the preventive strike. In Model 2, asymmetric information and the use of Bayesian learning are introduced. This is a crucial element, as the decision of the preventer often depends on the uncertain timeline of the breakout. In Model 3, the model is adjusted to include a multi-stage, stochastic approach, which mirrors the multi-stage technological hurdles of a nuclear program. Such an approach is crucial, as the effectiveness of intervention varies depending on the stage of the program (International Atomic Energy Agency, 2021). Lastly, Model 4 treats the proliferator's expenditure on deterrent capabilities as endogenous, which means that the cost of a strike becomes a function of the proliferator's decisions—a key component of traditional deterrence theory (Schelling, 1966).

The main contributions of the paper are as follows: First and foremost, the paper offers a mathematically precise and unified framework for analysing proliferation and pre-emption, as well as synthesizing different strands of thought into a coherent framework. Second, through the analysis of each of these models, the paper is also able to derive important strategic insights, such as the existence of a "pre-emption trap" and the circumstances under which "signalling" or "pooling" equilibria might occur. Third, the paper also offers a discussion of numerical solution methods for each of these models, which should be useful for policymakers. This paper formulates and analyses each of the models as well as presenting the key equations and analytical results for each model.

1.1. Ethical Considerations in Nuclear Proliferation Games"

The modeling of nuclear proliferation and preventive war does raise profound ethical questions that goes beyond mathematical formalation. These considerations Impacts on not only policy decisions but also the construction and interpretation of game-theoretic models.

A. Just War Theory and Preventive War

The ethical framework of jus ad bellum (justice of war) and jus in bello (justice in war) gives a lens through which to evaluate preventive strikes. Key considerations include:

  • Proportionality: Preventive strikes, as modeled in this paper, involve weighing the costs of intervention against the costs of inaction. The strike threshold x* derived in Theorem 1 implicitly encapsulates a proportionality judgment at what level of nuclear capability does the threat justify the costs of war?
  • Last Resort: The model's optimal stopping formulation captures the tension between acting early and exhausting peaceful alternatives. The continuation region represents a commitment to non-military alternatives, while the strike region represents the point at which these options are deemed exhausted.
  • Discrimination: Our multi-stage model (Model 3) does emphasize how different stages of a nuclear program have different features research facilities may have civilian applications, resulting to difficulty in the ethical assessment of targeting decisions.

B. Civilian Casualties and Collateral Damage

While our model treats strike cost S as a scalar, this parameter implicitly incorporates multiple ethical and material considerations:

  • Direct civilian casualties: Military strikes on nuclear facilities may cause immediate loss of life among facility workers and surrounding communities.
  • Environmental consequences: Attacks on nuclear installations risk radioactive contamination, potentially affecting populations across borders.
  • Indirect humanitarian impacts: Sanctions and sabotage (modeled through γ) may affect civilian populations through economic hardship, medical supply restrictions, and infrastructure damage.

C. The Ethical Asymmetry of Information

Model 2 incorporates asymmetric information about the proliferator's intentions. This information asymmetry creates an ethical dilemma:

  • Type I error (false positive): Mistakenly identifying a peaceful nuclear program as a weapons program could lead to unjust preventive war.
  • Type II error (false negative): Failing to identify a genuine weapons program risks the emergence of a nuclear-armed state with potential aggressive intent.

The Bayesian learning framework in our model formalizes this ethical tension, the preventer must balance the risks of both error types, each carrying different moral weights.

D. Moral Hazard and Commitment

Our model's treatment of deterrence (Model 4) raises ethical questions about strategic commitments:

  • Extended deterrence commitments: When a preventer commits to protecting allies, this forms moral obligations that may enforce an action even when direct national interests are not threatened.
  • Deterrence stability: The "preemption trap" identified in Theorem 4 suggests that certain paths lead inevitably to conflict. Recognizing these paths requires ethical reflection on how to structure the game to avoid such unavoidable conflict.

E. Justice in the Distribution of Risk

The proliferator and preventer in our model are treated as rational utility-maximizers. However, the true costs of conflict are distributed unevenly:

  • Civilian populations in both states bear risks disproportionate to their decision-making power
  • Regional neighbors may suffer from conflict externalities (refugee flows, environmental damage)
  • Future generations inherit the consequences of decisions made today

The model's discount rate r implicitly weighs present against future costs, representing an ethical judgment about temporal discounting.

F. Limitations of Our Ethical Framework

While we acknowledge that our mathematical model cannot completely capture all ethical considerations, the game-theoretic approach essentially abstracts away from individual human suffering, the moral weightiness of intentions, and the complexities of global legal frameworks. However, by making these trade-offs explicit, the model provides a structured framework within which ethical deliberations can occur.

G. Policy Implications of Ethical Considerations

The incorporation of ethical considerations produces several policy-relevant perceptions:

  1. The ethical case for early intervention is reinforced by the observation (Figure 6) that strike success probability drops below 30% at later stages, tardy action may increase civilian casualties on both sides.
  2. The value of intelligence (Figure 5) is not simply strategic but ethical, better information reduces the probability of both Type I and Type II errors.
  3. Deterrence investments (Figure 7) are ethically complex, they may stabilize the system by raising costs but also enable breakout, potentially increasing long-term risk.

2. Background Research

The strategic interaction between nuclear proliferators and potential preventers has long held a central place in the literature of international relations since the beginning of the nuclear age (Schelling, 1960; Sagan & Waltz, 2013). To understand the conditions under which states engage in nuclear proliferation and the conditions under which adversaries engage in preventive military action against proliferators, one must integrate the contributions of game theory, political science, and strategic studies. In this section, a summary of the major literature that undergirds the differential game approach presented here is offered.

2.1. Theories of Nuclear Proliferation

The motivations behind nuclear proliferation decisions have been extensively analysed from a variety of theoretical perspectives. In the security model of nuclear proliferation, states engage in nuclear proliferation decisions in response to existential security threats (Sagan & Waltz, 2013). In the domestic politics model of nuclear proliferation, the influence of bureaucratic interests and political coalitions is emphasized (Sagan, 1996). In the norms model of nuclear proliferation, the impact of international non-proliferation regimes and the nuclear taboos of states is emphasized (Singh & Way, 2004). In addition to these models of nuclear proliferation, quantitative studies of nuclear proliferation correlates are also prominent. In a quantitative analysis of nuclear proliferation decisions, Singh & Way (2004) found that states under significant security threats, with high industrial capabilities, and without security guarantees from nuclear-armed allies are likely to engage in nuclear proliferation decisions. Recent studies by Narang (2022) introduce the concept of proliferation rings to illustrate how nuclear proliferation occurs through a network of covert activity.

Nevertheless, most models of nuclear proliferation assume nuclear acquisition as a singular choice rather than a continuous one. This overlooks the continuous nature of nuclear acquisition, as it occurs through distinct stages, each with its own implications for nuclear strategy (International Atomic Energy Agency, 2021).

2.2. Preventive War and the Proliferation-Prevention Dilemma

The choice of whether to launch a preventive attack on a nascent nuclear program is one of the most important decisions a state must make. History has provided valuable lessons on this matter. The Israeli strike on Iraq’s Osirak nuclear reactor in 1981 was successful, but it came at a substantial diplomatic cost (Feldman, 2011). The Israeli strike on Syria’s Al Kibar nuclear plant in 2007 was successful with minute cost, but its implications are controversial (Kroenig, 2018).

The theoretical literature on preventive wars depicts the importance of power shifts and obligation problems. Powell (1990) showed that preventive wars become more likely when there is a significant and irreversible shift in the power balance, which the weaker state cannot credibly commit to accepting. More recent theoretical work by Bas and Coe (2016) builds upon this framework by introducing dynamic signalling, which suggests that a state may deliberately take its time in developing its proliferation capabilities to signal its benign intentions and avoid preventive wars. In this regard, there is a "proliferation puzzle" in which faster proliferation does not equate to more malevolent intentions. Debs and Monteiro (2014) introduced the concept of "known unknowns" in the proliferation domain, which suggests that the level of uncertainty over the proliferator's intentions and timeline is crucial in determining the decision calculus of the potential preventive attacker.

2.3. Game-Theoretic Approaches to Conflict and Proliferation

Game theory offers a powerful framework for understanding strategic interactions in international conflict. Initial attempts by Bueno de Mesquita and Riker (1982) to use game theory to analyse nuclear proliferation employed a two-stage approach but did not consider the dynamic nature of the process. With the development of differential game theory, a new tool was created for understanding dynamic strategic interactions. Simaan and Cruz (1975) are considered one of the first attempts to formalize the model of an arms race as a differential game. In this model, countries’ investments in the military are considered a dynamic process where one country reacts to the investments of the other. Later, Brito and Intriligator (1985) built on this model to consider conflict from a differential game perspective and demonstrated how the process of accumulating arms can lead to a Pareto inferior equilibrium from a cooperative perspective.

More recent attempts by Acemoglu and Wolitzky (2014) to model conflict and appropriation using a dynamic game approach highlighted the impact of future conflict on the decisions made by countries today. Although not exclusively focused on nuclear proliferation, this model offers a vital foundation for understanding how countries might allocate resources between economic and military investments.

2.4. Optimal Stopping and Real Options in Conflict Decisions

The decision to launch a preventive strike is formally analogous to the investment decision in the presence of uncertainty. Dixit and Pindyck (1994) pioneered the application of the real options approach. In the presence of uncertainty and irreversibility of investments, decision-makers have an incentive to delay investment decisions in the hopes of acquiring more information. This approach has been extended to the context of conflict decisions. Meirowitz and Sartori (2008) applied the optimal stopping problem to the context of war and demonstrated the potential for the option to delay war to reduce the probability of war, but only in the presence of precise information about relative capabilities. Baliga and Sjöström (2004) applied the real options approach to the context of arms races and demonstrated the paradoxical effect of ambiguity in intentions to stabilize deterrence.

The application of the real options approach to the context of nuclear proliferation is limited. Bas and Coe (2016) extended the dynamic signalling model by incorporating the elements of the optimal stopping problem. However, the continuous-time optimal stopping problem is not developed. This is the contribution of the real options approach proposed in Model 4.

2.5. Asymmetric Information and Signalling in Proliferation

The difficulty of accurately assessing the adversary's intentions and breakout timelines is a challenge in proliferation crises. Ferguson (2007) demonstrated the intelligence failures that have characterised the major proliferation crises, such as the underestimated Iraqi programme before 1991 and the overestimation of the Iraqi programme after 2003. Dynamic signalling models have addressed the question of how to signal information through the process of proliferation. Bas and Coe (2016) demonstrated the potential for slower proliferation to serve as a costly signal of benign intentions, thereby enabling the proliferator to establish a more credible deterrent without triggering preventive strikes. However, the effectiveness of the signalling mechanism relies on the ability of the preventer to accurately observe the investment decisions of the proliferator.

The effect of strategic ambiguity in proliferation has been studied by Baliga and Sjöström (2008), which concluded that strategic ambiguity can enhance deterrence by creating ambiguity over retaliatory capacity. But ambiguity can also lead to the risk of miscalculation, as the preventer might overestimate the capabilities of the proliferator and launch a preventive attack.

2.6. Multi-Stage Proliferation and Intervention Timing

The process of nuclear proliferation does not occur continuously but rather in distinct steps. Each step has unique characteristics with respect to observability and reversibility. The research and development phase comprises dual-use technology that makes it difficult to isolate civilian and military uses. The enrichment phase represents a serious point in the process. The capability to enrich uranium to weapons-grade stages brings the proliferator closer to weaponization. The weaponization/deployment phase makes it more problematic to intervene by giving the proliferator the ability to retaliate against an attack. The timing of military strikes against proliferating nations was examined by Fuhrmann and Kreps (2010), which found that strikes are most likely to occur in the enrichment phase. This matches with the theory of a "window of opportunity" that closes as the proliferating state develops further.

2.7. Deterrence Theory and the Stability-Instability Paradox

The theory of deterrence was first introduced by Schelling (1966), which focused on the importance of retaliation in the deterrence process. The stability-instability paradox argues that stable deterrence at the strategic level may lead to instability at lower levels. This occurs because nations are comfortable pursuing conventional wars due to the stability of their strategic nuclear forces. The problem with applying deterrence theory to emerging proliferators is the lack of second-strike capability. Zagare and Kilgour (2000) introduced the concept of "perfect deterrence." They found that stability in the deterrence process depends upon the credibility of retaliation. The credibility of retaliation depends upon the ability of the state to survive a first strike. This ability will be limited in emerging proliferators.

Narang (2015) has also identified different nuclear postures that are adopted by states during the process of nuclearization. These include the catalytic posture, assured retaliation posture, and asymmetric escalation posture. However, the process of dynamic progression of nuclearization by states from one posture to another has not been sufficiently researched.

2.8. Synthesis and Research Gap

The literature provides several valuable insights into the proliferation prevention problem. However, there are certain limitations of the literature that need to be addressed. Firstly, the literature has not sufficiently explored the dynamic progression of nuclearization by states from one posture to another. In addition, the literature has not sufficiently explored the interaction between the proliferator’s investment in nuclear capabilities and its investment in deterrent capabilities. In addition, the literature has not sufficiently explored the interaction between the proliferator’s investment in nuclear capabilities and its investment in deterrent capabilities. Furthermore, the literature has not sufficiently explored the multi-stage nature of nuclearization by states with stage-dependent observability and intervention effectiveness.

This paper addresses the above limitations of the literature by providing a comprehensive differential game framework that integrates:

  1. continuous state dynamics with an optimal stopping problem for preventive strike.
  2. asymmetric information and Bayesian learning about the proliferator’s type.
  3. the interaction between the proliferator’s investment in nuclear capabilities and its investment in deterrent capabilities.
  4. the multi-stage nature of nuclearization by states with stage-dependent observability and intervention effectiveness.
  5. multi-stage program progression with stage-dependent intervention effectiveness; and
  6. endogenous deterrence investment that directly affects the cost of preventive action.

By integrating all these factors into one mathematical structure, this research provides a foundation for the strategic analysis of nuclear proliferation and preventive war.

3. Mathematical Formulation

We present here four different, although interconnected, differential game models. They differ from the previous ones by the level of complexity in the strategies employed.

Model 1: The Basic Proliferation-Preemption Game

3.1.1 State Variable
Let x(t)[0,1] represent the proliferator's normalized nuclear weapons capability. The thresholds are defined as x=0 (no meaningful infrastructure), x=xB(0,1) (breakout threshold), and x=1 (fully operational nuclear weapons state).

3.1.2 Controls

  • Proliferator (Iran): Investment rate u(t)[0,uˉ]R+.
  • Preventer (Israel/US): Prevention intensity v(t)[0,vˉ]R+.

3.1.3 DynamicsThe state evolves according to:

x˙(t)=u(t)-γv(t)x(t)-δx(t),x(0)=x0[0,1] (1)where γ>0 is the effectiveness of prevention (e.g., sanctions, sabotage), and δ>0 is the natural decay rate of the nuclear program.

3.1.4 Preventive Strike Option
The preventer can launch a preventive strike at any time τ, causing an instantaneous reduction: x(τ+)=(1-κ)x(τ-), with κ(0,1). The strike incurs a lump-sum cost S>0. Multiple strikes are possible.

3.1.5 Objective FunctionalsBoth players minimize discounted costs over an infinite horizon with a common discount rate r>0.

  • Proliferator's Objective:

JI(u;v,τ)=E0e-rt12αu(t)2-βx(t)dt (2)where 12αu2 is the convex investment cost, and -βx is the benefit from nuclear capability.

  • Preventer's Objective:

JP(v,τ;u)=E0e-rt12ηv(t)2+θx(t)dtk=1Ne-rτkS (3)where 12ηv2 is the convex prevention cost, θx is the cost of the proliferator's capability, and the sum accounts for the lump-sum cost of N strikes.

3.1.6 Equilibrium Concepts
We consider two equilibrium concepts. In the Stackelberg equilibrium, the preventer acts as a leader committing to a strategy v(x) and strike policy τ(x), with the proliferator responding as a follower with u(x). In the Feedback Nash equilibrium, both players use Markov strategies u(x) and v(x).

3.1.7 Hamilton-Jacobi-Bellman (HJB) EquationsDefine value functions VI(x) and VP(x) for the continuation region x<xˉ.

Proliferator's HJB:

rVI(x)=minu[0,uˉ]12αu2-βx+VI'(x)(u-γv(x)x-δx) (4)The first-order condition yields u*(x)=-VI'(x)/α. Substituting gives:

rVI(x)=-12α[VI'(x)]2-βx-VI'(x)(γv(x)x+δx) (5)Preventer's HJB:

rVP(x)=minv[0,vˉ]12ηv2+θx+VP'(x)(u*(x)-γvx-δx) (6) The first-order condition yields v*(x)=γxηVP'(x).

3.1.8 Strike Boundary Conditions
At the strike threshold x=xˉ, the preventer is indifferent between striking and waiting, leading to:

  • Value matching: VP(xˉ)=VP((1-κ)xˉ)+S
  • Smooth pasting: VP'(xˉ)=(1-κ)VP'((1-κ)xˉ)

3.1.9 Linear-Quadratic SolutionFor the no-strike region, we conjecture quadratic value functions

 Vi(x)=12Aix2+Bix+Ci for i{I,P}.

Substitution leads to a Riccati equation for AP:

AP=γ2η(AP)2+2α(AP)2-2δAP-2rAP The positive solution yields the feedback strategies:

v*(x)=γηAPx+γηBP,u*(x)=-1α(AIx+BI) Model 2: Proliferation Game with Uncertain Breakout Capability

3.2.1 State, Type, and Information
The proliferator has a private type λ{λL,λH} with 0<λL<λH, representing its breakout speed. The state x(t) evolves as x˙(t)=λiu(t)-γv(t)x(t)-δx(t). The preventer observes x(t) with noise:

dy(t)=x(t)dt+σdW(t),y(0)=0 (7) where W(t) is a standard Wiener process. The preventer's belief is p(t)=P(λ=λHFty), with p(0)=p0.

3.2.2 Belief DynamicsThe Kushner-Stratonovich equation gives the evolution of the belief:

dp(t)=p(t)(1-p(t))(λH-λL)u(t)σ2dy(t)-x^(t)dt (8) where x^(t)=E[x(t)Fty]. In innovation form,

 dp=p(1-p)(λH-λL)uσdW~. (9)

3.2.3 Controls and Objectives
The proliferator's strategy is u(t)=ϕi(x^(t),p(t)) for type i. The preventer's strategy is v(t)=ψ(x^(t),p(t)). The objectives are analogous to Model 1, with the preventer's cost depending on x^.

3.2.4 HJB Equations in Sufficient Statistics
The problem reduces to states x^p. The proliferator's HJB for type H is:

rVIH=minu{12αu2-βx^+VIHx^λHu-γvx^-δx^+p(1-p)(λH-λL)2u2σ2+VIHpp(1-p)(λH-λL)uσCov(x,λ)σ+122VIHx^2Cov(x,λ)2σ2 (10)A similar equation holds for type L. The preventer's HJB is analogous but uses the expected controls and state.

3.2.5 Strike Boundary
The strike region S[0,1]×[0,1] satisfies a value matching condition that accounts for updated beliefs p' post-strike:

VP(x^,p)=VP((1-κ)x^,p')+S (11)

Model 3: Multi-Stage Proliferation with Escalating Tension

3.3.1 State Space and Transition Rates

Here, the state space is given by the finite set N={0,1,2,3},such that the state of proliferation is represented by n, i.e., 0 for R&D, 1 for enrichment, 2 for weaponization, and 3 for deployed weapons

The proliferator invests an amount un0, advancing with intensity qn,n+1(un)=un. The preventer chooses prevention vn0, causing degradation with intensity qn,n-1(vn)=γnvn, where γn is stage-dependent effectiveness.

3.3.2 Kolmogorov Forward EquationsLet πn(t)=P(n(t)=n). The dynamics are:

π˙0=-u0π0+γ1v1π1π˙1=u0π0-(u1+γ1v1)π1+γ2v2π2π˙2=u1π1-(u2+γ2v2)π2+γ3v3π3π˙3=u2π2-γ3v3π3 (12)3.3.3 Objective Functionals

  • Proliferator:

 JI=E0e-rtn=0212αnun(t)2-B1n(t)=3dt (13)

  • Preventer:

 JP=E0e-rtn=0312ηnvn(t)2+C1n(t)=3dtke-rτkSn(τk) (14)

3.3.4 Bellman EquationsLet VnI and VnP be the value functions in stage n. For n{0,1,2}:

  • Proliferator:

rVnI=minun012αnun2+un(Vn+1I-VnI)+γnvn(Vn-1I-VnI) (15)

with boundary V3I=0. The optimal control is un*=max0VnI-Vn+1Iαn.

  • Preventer:

rVnP=minvn012ηnvn2+C1n=3+un(Vn+1P-VnP)+γnvn(Vn-1P-VnP) (16)The optimal control is vn*=max0γn(VnP-Vn-1P)ηn.

Model 4: Preemption as a Real Option with Deterrence

3.4.1 State Variables
The state is now two-dimensional: x(t)[0,1] (nuclear capability) and d(t)[0,dˉ] (deterrence capability, e.g., conventional missiles, hardened facilities).

3.4.2 Dynamics

x˙(t)=u(t)-γv(t)x(t)-δxx(t),x(0)=x0d˙(t)=w(t)-δdd(t),d(0)=d0 (17)where w(t) is the proliferator's investment in deterrence.

3.4.3 Endogenous Strike Cost
The cost of the preventive strike is defined by the level of deterrence capability of the proliferator, which is given by the function S(d)=S0+ϕd, where the marginal cost of overcoming deterrence is greater than zero, i.e., ϕ>0

3.4.4 Objective Functionals

  • Proliferator:

JI=E0e-rt12αu(t)2+12ηw(t)2-βx(t)-ζd(t)dt (18)Preventer:

JP=E0e-rt12γvv(t)2+θx(t)dt+e-rτS(d(τ)) (19)

3.4.5 HJB Equations in Continuation Region
Define the continuation region C={(x,d):x<xˉ(d)}. The HJB for the proliferator is:

rVI=minu,w012αu2+12ηw2-βx-ζd+VIx(u-γvx-δxx)+VId(w-δdd) (20)First-order conditions give u*=-1αVIx and w*=-1ηVId.
The preventer's HJB is:

rVP=minv012γvv2+θx+VPx(u*-γvx-δxx)+VPd(w*-δdd) (21)with optimal prevention v*=γxγvVPx.

3.4.6 Strike Boundary ConditionsThe strike boundary Γ={(x,d):x=xˉ(d)} satisfies:

  • Value matching: VP(x,d)=VP((1-κ)x,d)+S0+ϕd
  • Smooth pasting:

VPxx,d=1-κVPx1-κx,d,

VPd(x,d)=VPd((1-κ)x,d)+ϕ

3.5. Formal Theorems

Theorem 1 (Existence of Strike Threshold)

Under standard convexity and discounting assumptions, there exists a unique threshold x* such that:

  • Strike is optimal if xx*
  • Continuation is optimal if x<x*

Proof.

We start by analyzing the infinite horizon stochastic control problem faced by the prevention state. Let xtR+ the nuclear capability of the proliferator at time t, follow the controlled stochastic differential equation:

dxt=(αut-βvt-δxt)dt+σdWt (22)

where ut denotes the investment of the proliferator, v t denotes the counter-proliferation effort of the preventer, and Wt is a standard Brownian motion representing the uncertainty. The decision of the preventer is defined by the stopping time  τ , which represents the time of the preventive strike, which involves a constant cost C > 0. The value function is:

V(x)=supτE0τe-rtπ(xt)dt-e-rτS1τ (23)

where π(x) is the flow payoff and r>0 is the discount rate.

The Hamilton-Jacobi-Bellman (HJB) equation for the continuation region S={x:V(x)>-S} is:

rV(x)=π(x)+LV(x)where L is the infinitesimal generator:

LV(x)=(αu*(x)-βv*(x)-δx)V'(x)+σ22V''(x) (24)For the stopping region S={x:V(x)=-S}, we have the value-matching condition:

V(x)=-Sand the smooth-pasting condition at the boundary x*:

V'(x*)=0Step 1: Existence. Define the operator:

TV(x)=supτE0τe-rtπ(xt)dt-e-rτS (25)

By standard results in optimal stopping theory (Øksendal, 2003), under the assumptions that π() is continuous, bounded, and satisfies appropriate growth conditions, and that xt is a regular diffusion, T is a contraction mapping on the space of continuous functions. Hence, a fixed point exists, giving a unique value function V.

Step 2: Concavity. The value function V inherits concavity from the convexity of costs. Specifically, for any x1,x2 and λ[0,1],

V(λx1+(1-λ)x2)λV(x1)+(1-λ)V(x2) (26)This follows because the payoff function π is concave and the stopping rule can be randomized. Concavity implies that the continuation region C is convex and the stopping region S is convex.

Step 3: Verification and Uniqueness. Define the candidate threshold:

x*=inf{x0:V(x)=-S} (27)By continuity of V and the fact that limxV(x)=- (since the flow payoffs are bounded and eventual strike is inevitable), such an x* exists. For x<x*V(x)>-S, so continuation is optimal. For x>x*V(x)<-S, so immediate strike is optimal.

Uniqueness follows from the strict concavity of V. If there were two distinct thresholds x1*<x2*, then for x(x1*,x2*) we would have both V(x)>-S (since x<x2*) and V(x)<-S (since x>x1*), a contradiction.

The smooth-pasting condition V'(x*)=0 provides the boundary condition that determines x* uniquely. This completes the proof.

Theorem 2 (Comparative Statics of Threshold)

The strike threshold x* satisfies:

  1. x*β<0 (more effective prevention lowers the threshold)
  2. x*S>0 (higher strike cost delays the strike)

Proof.

Let V(x;θ) denote the value function parameterized by θ{β,S}. At the optimal threshold x*(θ), the value-matching and smooth-pasting conditions hold:

V(x*;θ)=-S,Vx(x*;θ)=0Differentiate the value-matching condition with respect to θ:

Vxx*θ+Vθ=-SθSince Vx(x*;θ)=0 (smooth pasting), this simplifies to:

Vθ=-SθCase 1: θ=β (prevention effectiveness). The value function satisfies the HJB equation:

rV(x)=π(x)+(αu*(x)-βv*(x)-δx)Vx(x)+σ22Vxx(x) (28)Differentiating with respect to β and evaluating at x=x*:

rVβ=-v*(x*)Vx(x*)+σ22VxxβSince Vx(x*)=0, and by standard comparative statics results for elliptic operators (Peskir & Shiryaev, 2006), the term σ22Vxxβ has sign opposite to Vβ. This yields:

Vβ<0Now apply the differentiated value-matching condition. For βSβ=0, so:

Vβ=0This appears contradictory, but we must account for the dependence of x* on β through the smooth-pasting condition. The correct approach uses the implicit function theorem on the smooth-pasting condition Vx(x*;β)=0:

2Vx2x*β+2Vxβ=0Hence:

x*β=-2Vxβ2Vx2 (29)By concavity of V2Vx2<0. Standard envelope arguments give 2Vxβ<0 because higher β reduces the marginal benefit of waiting. Therefore, x*β<0.

Case 2: θ=S (strike cost). Differentiate the value-matching condition directly:

Vxx*S+VS=-1 (30)Since Vx(x*)=0 and VS=-1 (a direct consequence of the value function being linear in S), we obtain:

0x*S+(-1)=-1This holds identically, so we must differentiate the smooth-pasting condition instead. Differentiating Vx(x*;S)=0 with respect to S:

2Vx2x*S+2VxS=0Thus:

x*S=-2VxS2Vx2 (31)Now, 2VxS>0 because increasing the strike cost makes the stopping region smaller, shifting the threshold rightward. Since 2Vx2<0 by concavity, we conclude x*S>0.

Theorem 3 (Signaling Equilibrium)

A separating equilibrium exists iff VH(x,p)-VL(x,p)>0, where VH and VL are the value functions for high-type and low-type proliferators, respectively.

Proof.

Consider a dynamic game of incomplete information. The proliferator has private type θ{L,H}, where H represents a "breakout" type with higher investment efficiency αH>αL. The preventer holds a prior belief p0=P(θ=H) and updates beliefs via Bayes' rule.

Let Vθ(x,p) denote the value function of type θ when the preventer's belief is p. A separating equilibrium requires that:

  1. Incentive compatibility (IC) for the low type:

VL(x,p)VL(x,p=0)The low type prefers to reveal itself rather than mimic the high type.

  1. Incentive compatibility (IC) for the high type:

VH(x,p=1)VH(x,p)The high type prefers to be revealed as high rather than be mistaken for low.

  1. Belief consistency: The preventer's posterior belief pt evolves according to:

dpt=pt(1-pt)σdYt-μLdtσ (32)where Yt is the observed signal of capability, and μθ is the drift under type θ.

Step 1: Necessity. Suppose separating equilibrium exists. Then now of separation, the high type's action reveals its type. The payoff from separation must exceed the payoff from pooling, otherwise the high type would deviate. The difference VH(x,p)-VL(x,p) represents the informational rent or cost of signaling. For separation to be credible, the high type must find it worthwhile revealing:

VH(x,p)>VL(x,p)If this inequality were reversed, the high type would have an incentive to pool, contradicting separation.

Step 2: Sufficiency. Assume VH(x,p)-VL(x,p)>0. Construct the following separating strategy:

  • Low type: Choose investment uL(x) that maximizes its value given that the preventer will correctly infer its type. This yields VL(x,p=0).
  • High type: Choose investment uH(x) that reveals its type, possibly at some cost, but yields the separating payoff VH(x,p=1).

Define the belief update rule:

pt+=1if observed investment utuH(xt)ptif observed investment ut(uL(xt),uH(xt))0if observed investment utuL(xt) (33)

This is a "monotone" strategy where higher investment signals higher type.

Verification of incentive compatibility:

  • For the low type: By construction, choosing any uuH(x) would lead the preventer to believe it is high type, yielding continuation value VL(x,p=1). Since VL(x,p=1)<VL(x,p=0) (by the monotonicity of value functions in beliefs), deviation is suboptimal. Choosing u(uL,uH) leads to unchanged belief but lower current payoff, also suboptimal.
  • For the high type: Choosing u<uH(x) would lead the preventer to believe it is low type, yielding VH(x,p=0). Since VH(x,p=0)<VH(x,p=1) (again by monotonicity), deviation is suboptimal.

Step 3: the existence of equilibrium, as per standard results on dynamic signaling games (Mailath & von Thadden, 2013), the above-described monotone strategy profile would form a perfect Bayesian equilibrium under the condition that the single-crossing condition holds. It is apparent that the single-crossing condition holds since the marginal benefit of investment is increasing with type:  ux˙=αθ, with αH>αL. Thus, high types have a higher marginal return on investment.

The condition VH(x,p)-VL(x,p)>0 ensures that the high type's payoff from separation exceeds its payoff from pooling, making the separating equilibrium sustainable.

 

Theorem 4 (Preemption Trap)

There exists a manifold M in the state space such that:

  • Below M, the system follows a path of peaceful accumulation.
  • Above M, the system is driven to an inevitable strike.

Proof.

Consider the two-dimensional dynamical system with state variables xd, where x is nuclear capability and d is deterrence capability. The dynamics are given by:

x˙=f(x,d)=αu*(x,d)-βv*(x,d)-δxd˙=g(x,d)=γ(x,d)-ηd (34)where γ(x,d) represents the endogenous deterrence investment, and η>0 is the depreciation rate of deterrence.

Step 1: Equilibrium analysis. Define the nullclines:

x˙=0d=h1(x)d˙=0d=h2(x)Assume h1 is increasing and h2 is decreasing (or vice versa), ensuring a unique interior equilibrium x*d*.

Linearize the system around x*d*:

x˙d˙=Jx-x*d-d*+O((x,d)-(x*,d*)2) (35)where the Jacobian matrix is:

J=fxfdgxgd (36)Step 2: Saddle-point characterization. Suppose the determinant det(J)<0. Then the eigenvalues λ1 and λ2 satisfy λ1λ2=det(J)<0, so the eigenvalues are real and of opposite signs. Hence, x*d* is a saddle point.

Let λs<0<λu denote the stable and unstable eigenvalues, respectively. The stable manifold M is defined as:

M=(x,d)R2:limtϕt(x,d)=(x*,d*) (37)where ϕt is the flow of the dynamical system. By the stable manifold theorem (Hirsch, Pugh, & Shub, 1977), M is a one-dimensional C1 manifold tangent to the stable eigenvector at x*d*.

Step 3: Basin separation. The stable manifold M partitions the state space into two regions:

  • Region A={(x,d):trajectories converge to (0,0)} (peaceful equilibrium)
  • Region B={(x,d):trajectories diverge to  or hit strike boundary} (preemption trap)

By the properties of saddle-point dynamics, M is the boundary between these basins of attraction. For initial conditions below M (in the direction of the stable eigenvector), trajectories are attracted to the peaceful equilibrium. For initial conditions above M, trajectories are repelled away from the saddle point and eventually cross the strike threshold.

Step 4: Peaceful accumulation (below M). For xd with d<h2(x) (below the stable manifold), we have:

d˙<0andx˙<0for sufficiently small xThus, the system moves toward the origin 00, representing a stable equilibrium with low nuclear capability and low deterrence. The proliferator never reaches the strike threshold, so preventive war does not occur.

Step 5: Inevitable strike (above M). For xd with d>h2(x) (above the stable manifold), the dynamics exhibit:

d˙>0andx˙>0This creates a positive feedback loop: higher nuclear capability induces higher deterrence investment, which in turn justifies further nuclear buildup. The system evolves such that x(t) increases monotonically and crosses the strike threshold x* in finite time:

τ=inf{t0:x(t)x*}<At time τ, the preventer finds it optimal to strike, leading to conflict.

Step 6: Uniqueness of M. The stable manifold M is unique by the invariant manifold theorem. Any other manifold separating the basins would coincide with M by the uniqueness of the stable manifold in a neighborhood of the saddle point, and by global considerations of the flow.

Thus, the existence of such a manifold M is established, completing the proof.

4. Numerical Simulation

To complement the theoretical analysis and illustrate the dynamic behavior of the nuclear proliferation game, we conduct numerical simulations across all six figures. The simulations are implemented in Python using standard scientific computing libraries (NumPy, SciPy, Matplotlib). Below we detail the numerical methods employed for each class of simulations.

4.1 Convergence and Stability

All simulations satisfy standard convergence criteria:

  • Time step convergenceΔt=0.01 was verified to produce stable results with relative error < 10-4 compared to Δt=0.001
  • ODE integration: Absolute and relative tolerances set to 10-8 and 10-6, respectively
  • Monte Carlo: Belief evolution uses 1,000 independent paths (Figure 4) to ensure statistical stability

4.2 Computational Implementation

All simulations were performed in Python 3.10 with the following key libraries:

  • numpy (v1.24): Array operations and random number generation
  • scipy.integrate (v1.10): ODE integration
  • matplotlib (v3.7): Visualization and figure generation

4.2.1. State Evolution

The state variable x(t), representing nuclear capability, evolves according to the stochastic differential equation:

dxt=(αu(xt)-βv(xt)-δxt)dtWe discretize the time horizon 0T with step size Δt=0.01 using a forward Euler scheme:

xt+1=xt+(αu(xt)-βv(xt)-δxt)ΔtTwo scenarios are simulated:

  • Baseline: Investment function u(x)=0.3x, prevention effort v(x)=0.5x
  • High Investmentu(x)=0.6x, with prevention unchanged

The strike threshold is set at x*=0.7 based on theoretical predictions.

4.2.2. Strike Threshold vs. Cost

The strike threshold x* is derived from the optimal stopping problem solution. Following Theorem 2, we parameterize:

x*(S)=kS+cwhere k=0.4 and c=0.2 for baseline parameters. The alternative curve with β=1.2 uses k=0.35c=0.15. The function is evaluated at 50 evenly spaced points over S[0.5,3.0].

4.2.3. Phase Diagram

The two-dimensional dynamical system is defined as:

x˙=α(0.5x)-β(0.8x)-δx+0.1dd˙=γx-0.3dwith parameters α=0.5β=0.8δ=0.1γ=0.4. The vector field is computed on a 20×20 grid over (x,d)[0,1.2]2. To avoid division-by-zero errors in quiver plots, we apply a mask where the vector magnitude exceeds 10-10.

Trajectories are integrated using SciPy's odeint solver (LSODA algorithm) over t[0,30]. Initial conditions are:

  • Stable region: x0,d0)=(0.15,0.05
  • Unstable region: x0,d0)=(0.75,0.65

The stable and unstable manifolds are approximated by linear fits through the saddle point x*,d*)=(0.5,0.3 with slopes determined from the eigenvectors of the Jacobian.

4.2.4. Belief Evolution

The preventer's posterior belief evolves via Bayesian filtering. The observation process is:

dYt=μθdt+σdWtwith μH=0.15 for breakout type and μL=0.05 for non-breakout. We simulate two noise levels: σlow=0.1 and σhigh=0.4.

The belief update follows:

pt=pt-1lHpt-1lH+(1-pt-1)lL

where lθ=exp-12ΔYt-μθΔtσΔt2.

Simulations run over T=20 with Δt=0.05, using 4,000-time steps. The random seed is fixed at 42 for reproducibility.

4.2.5. Multi-Stage Transition

The nuclear program is modeled as a discrete-time Markov chain with four stages: Research, Enrichment, Weaponization, and Deployment. The transition matrix P is specified in Table 3.

Strike success probabilities are derived from a logistic function:

Pr(success)=11+ek(s-s0)with k=-3 and s0=2.5 (where s indexes stages 0–3). The resulting values are 0.90.70.30.05.

The Markov chain is visualized using a force-directed layout with node positions optimized for clarity. Transition probabilities are annotated along curved arcs to avoid overlapping.

4.2.6. 2D Strike Boundary

The strike boundary xˉ(d) is specified as a linear function:

xˉ(d)=ad+bwith a=0.6b=0.2 for baseline, and a=0.9 for the strong deterrence scenario. The boundary is evaluated at 100 evenly spaced points over d[0,1.2].

Endogenous strike cost is modeled as:

S(d)=1+2dreflecting the increasing difficulty of striking a more deterrence-capable adversary.

5. Results and Discussion

The analytical and numerical solutions to the four models yield several key strategic insights. The strategies are derived from the linear-quadratic solution with parameters α=0.5β=0.8γ=1η=0.3θ=5δ=0.1, ρ=0.05, r=0.05. Below we present the results from each modeling framework, followed by a synthesis of policy-relevant findings.

5.1 Model 1: Basic Strategic Dynamics

The feedback strategies for the proliferator and preventer as functions of the nuclear capability state x are presented in Figure 1.

phpgoJHem_esqrnh.png
Figure 1: Optimal Policies vs. Nuclear Capability

 

The discontinuity represented by this figure is indicative of a policy shift, and it is one that is purposeful and economically significant. This arises from the fact that there are two distinct behavioral regimes for the preventer:

  • Continuation region (x<0.7): In this regime, the preventer actively counters proliferation, with optimal prevention effort increasing in the proliferator's capability. The function v*(x)=3x+0.2 characterizes the optimal prevention effort prior to any strike.
  • Post-strike region (x0.7): Following a preventive strike, the preventer's effort drops sharply to a low constant (v*(x)=0.1). This reflects a post-conflict environment in which active counter-proliferation efforts are substantially reduced.

The jump at x*=0.7 marks the strike threshold x*. It is at this level of nuclear capability that the preventer changes from a containment policy to a pre-emptive strike against the proliferator. After the strike occurred, the level of threat perceived by the opponent drops dramatically, leading to a discrete jump in the level of preventive efforts. This jump captures the essence of the significant shift in the overall strategic situation that accompanies the proliferator's nuclear capability passing the threshold x*. This modelling approach is consistent with the result of Theorem 1, which guarantees the existence of a unique strike threshold x* at which the optimal policies of the two players discontinuously change from continuation to strike.

In the continuation region, the strategies of the two players are strictly increasing functions of x. This means that as the nuclear program progresses, the proliferator invests more to achieve nuclear breakout, and the preventer also increases its level of prevention. At the strike threshold  xˉ=0.7, a strike occurs, reducing the nuclear capability of the proliferator to (1-κ)x*=0.56. There is an immediate discrete reduction in the strategies of the two players. Figure 2 illustrates the dynamic process of the nuclear capability of the proliferator over time under different investment strategies. For the baseline scenario (blue solid line), the investment of the proliferator is represented by the equation u(x)=0.3x, while the counter-proliferation action of the preventer is represented by the equation v(x)=0.5x. In this case, the system converges to a stable point where nuclear capability starts at a certain point but converges to a low level of about 0.2. This is a case where the system is effectively prevented from attaining the strike threshold x*=0.7.

The high investment scenario (red dashed line), where the investment of the proliferator is represented by the equation u(x)=0.6x,, shows a case of runaway escalation. In this case, nuclear capability is constantly on the increase and crosses the strike threshold at around t=38. After this point, a preventive strike is made by the preventer. Therefore, there is a bifurcation of the system into a stable point and a point where the system is not stable.

The strike threshold is represented by a horizontal dashed line at x*=0.7. It is clear from the figure that the baseline scenario is always well below the strike threshold while the high investment scenario crosses the threshold and thus leads to conflict.

Figure 2
Figure 2: State Trajectory of Nuclear Capability

Figure 3 illustrates the relationship between the strike threshold x* and the fixed cost C associated with the initiation of a preventive strike. The upward-sloping curves capture the fundamental economic intuition of rising intervention costs allowing for a higher nuclear capability in the proliferator state and thus prolonging the time to initiate a preventive strike. The baseline curve is associated with the effectiveness of the preventive strategy β=0.8. The strike threshold increases from x*=0.48 at S=0.5 to x*=0.89 at S=3.0. The non-linear relationship is given by the function x*(S)=0.4S+0.2 which captures the convexity of the optimal stopping problem and reflects the decreasing marginal impact of costs on the strike threshold as costs become extremely high.

The alternative curve corresponds to an improvement in the effectiveness of the preventive strategy (β=1.2). In this scenario, the strike threshold is universally lower across all values of (S), as highlighted in Theorem 2 and given by x*/β<0. Figure 3 provides critical insights for policymakers and demonstrates the interplay between investments in counter-proliferation efforts β and the elevation of the perceived costs of the conflict to the adversary S, which drive the strike thresholds in opposing directions. Figure 4 illustrates the phase diagram of the two-dimensional dynamical system defined on the state space x. The vector field indicates the direction of motion over time, and the blue solid line indicates the stable manifold M, the critical boundary between two qualitatively distinct outcomes.

Figure 3
Figure 3: Strike Threshold vs. Strike Cost
Figure 4
Figure 4: Phase Diagram with Stable Manifold

Some of the salient features are: The origin 00 is a state of peaceful equilibrium. The trajectories originating from the lower-left region of the plane converge to the equilibrium point, which implies that the system tends to achieve stability when both nuclear and deterrence capabilities are low. The green trajectory is an example of a peaceful accumulation of capabilities, where the system asymptotically converges to the origin.

Secondly, the point (x*,d*)=(0.5,0.3) is a saddle point and corresponds to an unstable equilibrium point. The stable manifold M intersects this point, and its direction is given by the eigenvectors of the linearized system. The red dashed line indicates the unstable manifold.

Thirdly, the system trajectories that are initiated from above the stable manifold (such as the orange trajectory) show diverging behavior and move towards the upper right of space. This is the preemption trap of Theorem 4. After passing the separatrix, nuclear capability and deterrence capabilities grew without bound, leading inevitably to a preventive strike. The positive feedback mechanism is obvious from the vector field. Indeed, a higher x means a higher d, which warrants a higher x.

The phase diagram provides a geometric picture of the strategic landscape. It shows that the outcome of peaceful coexistence or certain conflict depends not on the absolute levels of capability but on their relative position relative to the stable manifold. This highlights the need for early intervention, as once the system enters the preemption trap region, no other action than a preemptive strike can steer it towards peace.

The evolution of the posterior belief p(t)=P(θ=HFt) for the preventer is shown in Figure 5 for two noise conditions. The true type is the breakout type, i.e., θ=H, with a drift of μH=0.15 compared to the non-breakout type with a drift of μL=0.05.

In the presence of low levels of observation noise (σ=0.1, blue solid line), the belief process exhibits strong converging properties towards the true state. As can be seen, by approximately t10, the probability associated with the type of proliferator being a preventer exceeds 0.95. This allows for a strong inference of the type of proliferator and enables the appropriate policy responses. The smooth and monotonically increasing nature of the belief process reflects the high signal-to-noise ratio of the underlying observation process.

In the presence of high levels of observation noise σ=0.4, red dashed line), the belief process exhibits rather different behavior. As can be seen, the belief process exhibits high levels of volatility with the probability oscillating between 0.3 and 0.9 over long periods of time. Even at t=20, the belief process has yet to fully converge to the true state of the world. This high level of volatility poses a strategic problem to the preventer. On one hand, failing to act early due to high levels of noise may result in a Type I error. On the other hand, failing to act against a proliferator of type 'Breakout' results in a Type II error.

The gray dotted line at p=0.5 represents the prior belief, while the green dotted line at p=1 represents the true state. The growing divergence of the low-noise and high-noise trajectories over time captures the effect of compounding uncertainty: each new piece of information provides progressively less information in the high-noise case, thereby prolonging the time to convergence.

This plot directly confirms Theorem 3, which specifies the conditions for separating equilibriums. In the low-noise case, the high type has a credible way of signaling their types by investments; therefore, we achieve efficiency. In the high-noise case, however, the ability of the high types to signal their types diminishes, and we may have a pooling equilibrium where types are indistinguishable. The policy lesson is clear: the quality of intelligence is not just a tactical advantage but a strategic necessity for crisis stability.

Figure 5
Figure 5: Belief Evolution under Uncertainty

Figure 6 shows the progression of the nuclear program through distinct developmental stages, represented in two different visualizations. The left panel shows the structure of the Markov chain transitions, while the right panel shows the probability of strike success for each stage.

From the Markov chain transitions (left panel), it can be observed that the program development progresses through the following four stages: Research, Enrichment, Weaponization, and Deployment. The transition probabilities reveal several important characteristics. The research stage shows high probability of persistence P=0.6, indicating the long time required for initial research on nuclear development. The main transition probability from the research stage is to the enrichment stage P=0.35while the probability of direct transition to the next stage is low. The enrichment and weaponization stages show increasing probability of transition to the deployment stage (P=0.45 and P=0.60, respectively),

Figure 6
Figure 6: Multi-Stage Transition Probabilities

The right-hand panel illustrates the probability of success of a preventive strike against a nuclear program with the intent of neutralizing it, which is represented as a function of the nuclear program’s developmental stage. The probability of success of a preventive strike against a nuclear program shows a steep fall across the various stages of a nuclear program’s development. The probability of success of a preventive strike against a nuclear program begins at 90% in the Research stage but eventually reduces to 5% in the Deployment stage. The important point to note here is that there exists a critical point of concern between the Enrichment and Weaponization stages. The region of interest in the red area of the graph, which represents the Weaponization and Deployment stages of a nuclear program, represents the point of no return. At this point, the probability of success of a preventive strike against a nuclear program dip to less than 30%, and the cost of a preventive strike against a nuclear program becomes prohibitive.

This figure provides empirical guidance with respect to the timing of intervention. The successful implementation of preventative measures at the Research or early Enrichment stages offers the highest probability of success with the fewest potential side effects. The policy of waiting until there is definitive evidence of weaponization risks allowing the program to reach the point of no return, thus limiting the overall effectiveness of preventative measures and leaving policymakers with the unpalatable choice of living with a nuclear threat or launching an unsuccessful attack. The relationship between deterrence capability and the strike boundary is explored in Figure 7. This figure presents two panels. The left-hand panel presents the strike boundary xˉ(d) in two-dimensional state space, and the right-hand panel presents the endogenous relationship between deterrence and strike cost.

In the left panel, the strike boundary, indicated by the blue solid line, partitions the state space into two distinct areas. The first area, the Safe Region (green shaded area), is the region below the strike boundary where the preventer retains containment. The second area, the Strike Region (red shaded area), is the region above the strike boundary where immediate preventive action is optimal. The strictly increasing relationship of the strike boundary with respect to d is consistent with the idea that the cost of striking increases with the level of deterrence, thus allowing the proliferator to develop more nuclear capability before the preventive action.

The dashed red line represents an alternative scenario with stronger deterrence effectiveness (a=0.9 versus a=0.6 in the baseline). Under this scenario, the boundary shifts upward, expanding the safe region. This indicates that investments in deterrence, whether through hardened delivery systems, survivable second-strike capabilities, or extended deterrence commitments, can increase the proliferator's "breathing room" before the preventer finds it optimal to strike.

The right-hand panel explicitly illustrates the endogenous strike cost function S(d)=1+2d, which shows a linear relationship with respect to deterrence capabilities. The purple shaded region illustrates the cost premium resulting from deterrent investments. This relationship captures a critical aspect of the strategic trade-off: deterrence investments have the potential to deter a strike by increasing the cost of conflict while also allowing the proliferator to aggressively pursue nuclear capabilities, which could increase the probability of the system moving above the strike boundary.

Together, the two panels illustrate the two faces of deterrence in the proliferation’s prevention game: the potential for deterrence to be stabilized by increasing the conflict costs and expanding the region of peaceful accumulation, while also being destabilizing by allowing for breakout. This trade-off is captured in Theorem 4, where the preemption trap emerges when the deterrent investments push the system above the stable manifold.

Figure 7
Figure 7: 2D Strike Boundary and Endogenous Strike Cost

6. Sensitivity Analysis of Strike Threshold (x*)

The comparative static results for the optimal strike threshold, as derived in Theorem 2 and verified by numerical simulations, are summarized in Table 2. The results provide insight into how changes in some key parameters affect the preventive actor’s tolerance for nuclear capability accumulation before executing the preventive strike. Prevention Effectiveness β: An improvement in β, which represents the effectiveness of counter-proliferation technologies or strategies, lowers the optimal strike threshold x*/β<0. The seemingly counterintuitive effect of this variable can be attributed to the strategic commitment effect. In the face of very effective prevention, the preventive actor can intervene at a lower level of capability with a high probability of success. Thus, instead of waiting for the proliferator to accumulate more capability, the preventive actor strikes earlier, effectively limiting the latter’s scope for action.

Strike Cost S: Higher strike costs mean the decision threshold increases x*/S>0, which delays the decision to initiate the strike. This corresponds with the notion that costly interventions, measured by military, diplomatic, or economic metrics, should intuitively lead the preventer to allow the proliferator a greater level of nuclear capability before acting. The relationship is nonlinear, meaning diminishing marginal effects occur as the cost is extreme. Uncertainty σ2: Increased observation noise delays the decision x*/σ2>0, which corresponds with the notion of a ‘wait and see’ approach. In the face of noisy signals related to the true capability or intent of the proliferator, the decision to initiate the irreversible action of the strike is delayed in order to obtain more information. The effect of this, as shown in Figure 5, highlights the importance of information in the model, where greater information leads to earlier intervention.

Capability Decay δ: The accelerated natural decline in nuclear capability reduces the strike threshold x*/δ<0. If the proliferator’s capability is declining quickly due to technological obsolescence or other factors, the preventer may not feel compelled to act quickly to prevent nuclear breakout. The preventer may wait and see because the threat of nuclear breakout is likely to diminish over time. These comparative static effects define a framework for strategic policy formation. Those factors increasing the effectiveness of the preventer’s strategy β or the decay of the proliferator’s nuclear capability δ allow for earlier intervention thresholds. Those factors increasing the costs of striking S or the quality of the intelligence received σ2 raise the thresholds for intervention and thus increase the likelihood of the proliferator breaking out before intervention by the preventer.

Table 2: Sensitivity Analysis of Strike Threshold (x*)

Parameter

Effect on x*

Strategic Interpretation

β

x*

More effective prevention makes earlier intervention optimal

S

x*

More military expenditure makes it optimal to wait longer before striking

σ2

x*

More uncertainty makes it optimal to wait and see before acting

δ

x*

Faster decay of capabilities makes it optimal to strike sooner

7. Policy Implications

The theoretical and numerical results yield several key policy-relevant insights:

  1. Early Intervention Reduces the Probability of War: As illustrated in Figures 2 and 6, waiting until the evidence of the development of a nuclear weapon is undeniable allows the program to proceed beyond the point of no return, at which the probability of a successful strike is reduced to less than 30 percent. The preemptive actions taken during the research phase or the initial stages of enrichment have the highest probability of achieving the goal of neutralization with the least risk of escalation.
  2. Intelligence Uncertainty Delays Action

Figure 7 demonstrates the double role of deterrence; increasing the costs of strikes expands the safe region and stabilizes the system but also allows for the accumulation of nuclear weapons before a strike and is therefore destabilizing. It is necessary to manage deterrent investments to keep the system below the stable manifold of Figure 4.

  1. Deterrent Investments Create Strategic Trade-offs

Figure 7 demonstrates the double effect of deterrence: increasing strike costs makes the safe region larger but also allows for more nuclear power before a strike occurs. The policy-maker must design the deterrence investments such that the system operates below the stable manifold of Figure 4.

  1. Delayed Response Creates Preemption Traps

The phase diagram in Figure 4 illustrates that upon crossing the stable manifold M, trajectories diverge to conflict. This pre-emption trap implies that indecision in reaction to the development of a nuclear program may confine policymakers to a course of future conflict that might have been avoided by intervention at an earlier time.

8. Conclusion

The present work provides a comprehensive unified framework for analyzing the strategic interaction between nuclear proliferation and preventive war. It achieves this through the development of four progressively sophisticated types of differential game models that capture the key aspects of this complex strategic interaction: continuous development of capabilities, the threat of attack, asymmetric information, sequential development of technology, and endogenous deterrence formation.

The mathematical results clearly indicate that the strategic environment permits equilibriums, ranging across the spectrum from deterrence success to preemptive wars, depending critically upon the initial conditions, the structure of the information, and the cost of conflict. The “preemption trap” identified in Model 4 stands out as one of the most striking results, highlighting the potential for investments in deterrence to fuel conflict. The signaling dynamics in Model 2, too, indicate how the ambiguity of a proliferator’s activities might serve as a double-edged sword, potentially lengthening the time horizon of the crisis even as it increases the risk of miscalculations.

The analysis also indicates a path of future research. The model can be calibrated with the historical data available on nuclear programs, defense spending, and the effectiveness of sanctions. The numerical models also provide policymakers with a guide to compute the results of different interventions and engagements. In conclusion, this comprehensive model demonstrates that the proliferation problem is not a static choice but a dynamic interaction that is rich in information and strategic. The intricacies of the problem must be understood if one of the most intractable and dangerous challenges in global security is to be effectively addressed.

8.1. Assumptions, Limitations, and Significance"

A. Assumptions of the Study

The mathematical models developed in this paper rest on several foundational assumptions, which we explicitly state:

1. Rational Actor Assumptions:

  • Both proliferator and preventer are rational expected utility maximizers
  • Perfect rationality with common knowledge of rationality (except in Model 2 where information is asymmetric)
  • Both players have well-defined utility functions that are continuous and convex in costs

2. Economic Assumptions:

  • Investment costs are convex quadratic (12αu2), capturing diminishing returns
  • Capabilities follow continuous dynamics described by differential equations
  • All costs and benefits are measurable in comparable units (monetizable)
  • A common discount rate r applies to both players

3. Information Structure Assumptions:

  • Model 1: Complete information
  • Model 2: Asymmetric information with publicly observable state but private type
  • The observation process dy(t)=x(t)dt+σdW(t) correctly captures the signal structure
  • Beliefs follow Bayesian updating with the Kushner-Stratonovich equation

4. Strategic Interaction Assumptions:

  • Markov strategies (feedback strategies) depend only on current state variables
  • The Stackelberg and Feedback Nash equilibria are appropriate equilibrium concepts
  • Credible commitments can be made (Model 1 Stackelberg)
  • Preventive strikes, when implemented, are instant and irreversible

5. Modeling Simplifications:

  • Single proliferator and single preventer (two-player game)
  • The proliferator's type is binary (λH or λL)
  • Nuclear programs progress through four discrete stages
  • Strike success probability is monotonically decreasing in stage

B. Limitations

1. Theoretical Limitations:

  • Omission of Multi-Polar Dynamics: The model considers two players; real-world nuclear games involve multiple states (e.g., Iran-Israel-US-Iran, or regional competitions). The emergence of multiple nuclear powers creates more complex strategic interactions.
  • No Nuclear Terrorism Risk: The model does not consider the risk of nuclear weapons falling into non-state actors' hands, which significantly affects intervention decisions.
  • Abstracted Domestic Politics: Domestic political constraints on both players are not modeled. Bureaucratic incentives, leadership survival, and public opinion can override strategic calculations.
  • No Alliance Dynamics: Extended deterrence and alliance commitments are not formally modeled, though the Iran-US-Israel dynamic is referenced contextually.
  • Stationary Equilibrium Assumption: The infinite horizon models assume time-invariant parameters, while real-world parameters shift with technology, leadership, and international norms.

2. Empirical Limitations:

  • Calibration Challenges: Key parameters (αβγηθ) lack definitive empirical estimates, making policy predictions uncertain.
  • Data Scarcity: The small number of historical nuclear proliferation cases limits model validation.
  • Counterfactual Difficulty: We cannot observe what would have happened without intervention, making it hard to validate model predictions.
  • Case Selection Bias: Historical examples (Osirak, Syria, Iraq) may not be representative of all proliferation cases.

3. Modeling Limitations:

  • Linear-Quadratic Approximation: The LQ solution provides analytical tractability but may not capture all strategic nonlinearities.
  • Markov Assumption: Strategies depend only on current state, not on history or reputation.
  • No Bargaining Phase: The model does not include explicit negotiations or bargaining processes prior to conflict.
  • Continuous State Simplification: Nuclear capability is represented as a continuous variable, while real programs have discrete milestones.
  • Absence of Accidents: The model does not account for accidental nuclear detonations, which are a significant concern.

4. Computational Limitations:

  • Deterministic Simulations: Most simulations assume deterministic dynamics; stochastic simulations (beyond belief evolution) are limited.
  • Numerical Stability: Solutions depend on parameter choices; some parameter regions may exhibit numerical instability.
  • Limited Sensitivity Analysis: A complete global sensitivity analysis was beyond the scope of this paper.

C. Significance of the Study

Despite these assumptions and limitations, the paper makes several significant contributions to the literature and policy practice:

1. Theoretical Contributions:

  • Unified Framework: First comprehensive differential game model integrating continuous dynamics, optimal stopping, asymmetric information, multi-stage progression, and endogenous deterrence.
  • Novel Theorems: Formal proofs of strike threshold existence (Theorem 1), comparative statics (Theorem 2), signaling equilibrium conditions (Theorem 3), and the preemption trap (Theorem 4).
  • Mathematical Innovations: Extension of real options theory to nuclear proliferation; integration of Bayesian learning with differential games; stage-dependent Markov chain formulation.

2. Policy Significance:

  • Intervention Timing: Clear guidance on when intervention is most effective (research/enrichment stages, Figure 6), avoiding the "point of no return."
  • Intelligence Value: Demonstrates that intelligence quality (Figure 5) is a strategic necessity, not merely a tactical advantage.
  • Cost-Benefit Analysis: Provides a framework for evaluating the trade-off between strike costs and containment costs.
  • Deterrence Design: Identifies conditions under which deterrence investments are stabilizing versus destabilizing (the "two faces" of deterrence).

3. Strategic Insights:

  • Preemption Trap Identification: Formal demonstration that certain paths lead inevitably to conflict, highlighting the importance of early intervention.
  • Signaling Mechanism: Conditions under which a proliferator can credibly signal benign intentions through slow development.
  • Information Value: Quantifies the value of reducing uncertainty in proliferation crises.

4. Methodological Contributions:

  • Cross-Model Comparison: Progressive model development allows comparison of marginal contributions of each modeling element.
  • Analytical Tractability: LQ solutions provide benchmark cases for numerical analysis.
  • Computational Framework: Python-based implementation enables replication and extension.

5. Future Research Agenda:

  • Calibration: Using historical data to estimate parameters for specific proliferation cases.
  • Policy Counterfactuals: Evaluating alternative strategies through the model's lens.
  • Multi-State Extension: Expanding the framework to include multiple proliferators and preventers.
  • Bargaining Integration: Adding diplomatic negotiations as a state variable.
  • Stochastic Extensions: Incorporating full stochastic dynamics and option pricing methods.

6. Practical Applicability:

  • Policy Simulation: Policymakers can use the model structure to think systematically about proliferation scenarios.
  • Scenario Planning: The phase diagram (Figure 4) provides visual guidance on early warning signs.
  • Investment Prioritization: Helps identify which types of investments (intelligence, sanctions, deterrence) have the highest strategic returns.

9. Summary of Findings

  1. Model 1 (Basic Model): Establishes the foundational trade-off. A unique feedback Nash equilibrium exists where both players' strategies are monotonic in nuclear capability. The strike threshold is increasing in the strike cost S and decreasing in the preventer's cost of proliferation θ.
  2. Model 2 (Asymmetric Information): The presence of private information about breakout speed creates signaling incentives. A separating equilibrium can occur where the slow type invests less to mimic the fast type. The preventer's learning process via noisy observation can delay or accelerate the strike decision.
  3. Model 3 (Multi-Stage): The optimal timing of intervention is stage dependent. Preemption is most likely during the weaponization stage (Stage 2), as earlier stages are less threatening, and later stages (Stage 3) make the strike prohibitive due to assured retaliation.
  4. Model 4 (Endogenous Deterrence): The endogenous investment in deterrence creates a "preemption trap." When nuclear and deterrent capabilities are complementary, the system can exhibit a stable manifold that separates a peaceful equilibrium from a path leading to inevitable conflict. Deterrence can be stable if the proliferator can credibly invest in retaliation capabilities before crossing a critical nuclear threshold.

10. Statements

(a) Conflict of Interest: The author declares no conflict of interest regarding the publication of this manuscript. No financial or personal relationships have influenced the work reported in this paper.

(b) Funding: This research received no specific grant from any funding agency in the public, commercial, or non-profit sectors.

(c) Acknowledgments: The author would like to thank the anonymous reviewers for their valuable feedback on earlier versions of this work.

(d) Disclosure Statement (to be added to Acknowledgements):
The authors acknowledge the use of ChatGPT (OpenAI) for language polishing and grammatical refinement of the manuscript. All mathematical content, model development, theoretical analysis, numerical simulations, and interpretative conclusions are the sole work of the authors

Abbreviations

Abbreviation

Meaning

HJB

Hamilton-Jacobi-Bellman

IAEA

International Atomic Energy Agency

IC

Incentive Compatibility

ODE

Ordinary Differential Equation

R&D

Research and Development

SDE

Stochastic Differential Equation

References

  1. Acemoglu, D., & Wolitzky, A. (2014). The economics of conflict. Journal of Political Economy, *122*(6), 1215–1269. https://doi.org/10.1086/678995
  2. Baliga, S., & Sjöström, T. (2004). Arms races and negotiations. The Review of Economic Studies, *71*(2), 351–376. https://doi.org/10.1111/j.1467-937X.2004.00287.x
  3. Baliga, S., & Sjöström, T. (2008). Strategic ambiguity and arms proliferation. Journal of Political Economy, *116*(6), 1023–1057. https://doi.org/10.1086/595620
  4. Bas, M. A., & Coe, A. J. (2016). A dynamic theory of nuclear proliferation and preventive war. International Organization, *70*(4), 655–685. https://doi.org/10.1017/S0020818316000220
  5. Bas, M. A., & Schub, R. (2017). The predictive value of nuclear latency. International Studies Quarterly, *61*(3), 677–690. https://doi.org/10.1093/isq/sqx028
  6. Biddle, T. D. (2021). The specter of defeat: Intelligence and the decision for war. Princeton University Press.
  7. Brito, D. L., & Intriligator, M. D. (1985). Conflict, war, and redistribution. American Political Science Review, *79*(4), 943–957. https://doi.org/10.2307/1956244
  8. Bueno de Mesquita, B., & Riker, W. H. (1982). An assessment of the merits of selective nuclear proliferation. Journal of Conflict Resolution, *26*(2), 283–306. https://doi.org/10.1177/0022002782026002004
  9. Bunn, M., & Wier, A. (2006). Securing the bomb 2006. Project on Managing the Atom, Harvard University.
  10. Cirincione, J., Wolfsthal, J. B., & Rajkumar, M. (2010). Deadly arsenals: Nuclear, biological, and chemical threats. Carnegie Endowment for International Peace.
  11. Debs, A., & Monteiro, N. P. (2014). Known unknowns: Power shifts, uncertainty, and war. International Organization, *68*(1), 1–31. https://doi.org/10.1017/S0020818313000273
  12. Dixit, A. K., & Pindyck, R. S. (1994). Investment under uncertainty. Princeton University Press.
  13. Dockner, E. J., Jørgensen, S., Van Long, N., & Sorger, G. (2000). Differential games in economics and management science. Cambridge University Press.
  14. Feldman, S. (2011). Israeli nuclear deterrence: A strategy for the 1980s? Columbia University Press.
  15. Ferguson, C. D. (2007). The four faces of nuclear terrorism. Center for Nonproliferation Studies.
  16. Fudenberg, D., & Tirole, J. (1991). Game theory. MIT Press.
  17. Fuhrmann, M., & Kreps, S. E. (2010). Targeting nuclear programs in war and peace: A quantitative empirical analysis, 1941–2000. Journal of Conflict Resolution, *54*(6), 831–859. https://doi.org/10.1177/0022002710371689
  18. Gartzke, E., & Jo, D. J. (2014). Nuclear weapons and conflict: A theoretical and empirical re-assessment. Journal of Conflict Resolution, *58*(3), 403–429. https://doi.org/10.1177/0022002713515678
  19. International Atomic Energy Agency. (2021). IAEA safeguards: A historical perspective. IAEA.
  20. Kroenig, M. (2018). The logic of American nuclear strategy: Why strategic superiority matters. Oxford University Press.
  21. Mailath, G. J., & von Thadden, E. L. (2013). Incentive compatibility and the structure of institutions. Journal of Economic Theory, *148*(6), 2357–2380. https://doi.org/10.1016/j.jet.2013.08.001
  22. Meirowitz, A., & Sartori, A. E. (2008). Strategic uncertainty is a cause of war. Quarterly Journal of Political Science, *3*(4), 327–352. https://doi.org/10.1561/100.00007033
  23. Narang, V. (2015). Nuclear strategies and the challenge of nuclear proliferation. Annual Review of Political Science, *18*, 437–456. https://doi.org/10.1146/annurev-polisci-053013-040540
  24. Narang, V. (2022). Seeking the bomb: Strategies for nuclear proliferation. Princeton University Press.
  25. Parsi, T. (2017). Losing an enemy: Obama, Iran, and the triumph of diplomacy. Yale University Press.
  26. Powell, R. (1990). Nuclear deterrence theory: The search for credibility. Cambridge University Press.
  27. Powell, R. (2004). The inefficient use of power: Costly conflict with complete information. American Political Science Review, *98*(2), 231–241. https://doi.org/10.1017/S0003055404001116
  28. Sagan, S. D. (1996). Why do states build nuclear weapons? Three models in search of a bomb. International Security, *21*(3), 54–86. https://doi.org/10.2307/2539273
  29. Sagan, S. D., & Waltz, K. N. (2013). The spread of nuclear weapons: An enduring debate (3rd ed.). W. W. Norton & Company.
  30. Schelling, T. C. (1960). The strategy of conflict. Harvard University Press.
  31. Schelling, T. C. (1966). Arms and influence. Yale University Press.
  32. Simaan, M., & Cruz, J. B. (1975). Formulation of arms control and conflict problems as dynamic games. IEEE Transactions on Systems, Man, and Cybernetics, *5*(2), 225–232. https://doi.org/10.1109/TSMC.1975.5408402
  33. Singh, S., & Way, C. R. (2004). The correlations of nuclear proliferation: A quantitative test. Journal of Conflict Resolution, *48*(6), 859–885. https://doi.org/10.1177/0022002704269655
  34. Zagare, F. C., & Kilgour, D. M. (2000). Perfect deterrence. Cambridge University Press.
  35. Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of Dialogic Academic Presses (DAPresses), Dialogic Solutions Ltd and/or the editor(s). Dialogic Solutions Ltd and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

 

How to cite this article

Innocent Adinebo Katule and Akindele Michael Okedoye (2026). Dynamic Differential Games of Nuclear Proliferation and Preemption under Uncertainty. Dialogic STEM Journal, 1(1), 1-32. https://doi.org/10.66845/dstem.2026.00003

APA style shown. Use “Cite” for MLA, Chicago, IEEE, BibTeX and RIS formats.

Metrics

250

Article views

66

PDF downloads