Data Augmentation in Bayesian Estimation of the CIR Model
Summary
The document asks what augmented data means in Bayesian estimation of the Cox–Ingersoll–Ross (CIR) process. It describes estimating the model’s parameters from observations of a square-root diffusion and gives a discrete-time approximation. The paper’s notation distinguishes observed values from additional intermediate values inserted between observations, with multiple latent values associated with each interval.
The answer clarifies that data augmentation is a Bayesian inference technique, not a finance-specific concept. It introduces latent or missing data into the estimation procedure, making it possible to work with a more tractable conditional problem in iterative methods such as a Gibbs sampler. The response points to foundational reading on calculating posterior distributions by data augmentation, but does not explain the CIR-specific sampling steps, identify the full conditional distributions, or show how many intermediate points to add. Those implementation details must be obtained from the cited paper or a fuller treatment of the method.
Key ideas
- The CIR process models a state variable with mean-reverting drift and volatility proportional to its square root.
- Augmented data refers to additional latent observations inserted between the observed data points.
- Data augmentation is a Bayesian inference technique rather than a finance-specific concept.
- Adding latent data can make parameter estimation more tractable within iterative procedures such as Gibbs sampling.
- The discussion does not specify the CIR model’s full conditional distributions or sampling algorithm.
Tags
Full text
# What is augmented data when simulating stochastic differential equations using Gibbs Sampler?
# What is augmented data when simulating stochastic differential equations using Gibbs Sampler?
I am reading this paper on Bayesian Estimation of CIR Model.
Basically, it is about estimating parameters using Bayesian inference.
It estimates this stochastic differential equation: $$dy(t)=\{ \alpha-\beta y(t)\}dt+ \sigma \sqrt{y(t)}dB(t)$$ where $B(t)$ is standard Brownian motion.
by using this approximation: $$y(t+ {\Delta}^{+})=y(t)+\{\alpha-\beta y(t) \}{\Delta}^{+}+\sigma \sqrt{y(t)} {\epsilon}_{t}$$ ${\epsilon}_{t} \tilde{\ }N(0, {\Delta}^{+})$
My question is: Let $Y=({y}_{1},...,{y}_{T})$ denote observation data and ${Y}^{*}=(y_{1}^{*},...,y_{T-1}^{*})$ be AUGMENTED data, where $y_{*}^{t}=\{ y_{t,1}^{*},..., y_{t,M}^{*} \}$
What is augmented data?
I see that $y_{1}^{*}$ has elements $y_{1,1}^{*},..., y_{1,M}^{*}$. Is this what's so called augmented data? Why do we need this?
This seems like a finance concept I am missing.
## Answer by Alexey Kalmykov (score 4)
https://quant.stackexchange.com/a/3326
This is not a finance concept. Augmented data is related to Bayesian inference. It's essentially a way to improve maximum likelihood estimation from incomplete data. For details see the article "The calculation of posterior distributions by data augmentation" by Martin A. Tanner and Wing Hung Wong (it's referred in the paper you are reading).Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.