We all wish we had a crystal ball. For an importer-exporter, knowing the value of the Indian Rupee against the US Dollar next month could mean the difference between profit and loss. For a business, forecasting sales can drive inventory, hiring, and strategy. While we don’t have magic, we do have powerful statistical tools. One of the most robust and widely respected methods for this kind of crystal-gazing is ARIMA modelling. Itโ€™s a cornerstone of time series forecasting, and thanks to the R programming language, itโ€™s more accessible than ever. This guide will walk you through the entire process, from data to forecast, using R as our workshop.

Table of Contents

Getting to know our tools: R and the ‘forecast’ package

First, let’s meet our tools. R is a free, open-source software environment for statistical computing and graphics. It’s the language of choice for statisticians, data scientists, and academics because of its power, flexibility, and a massive community that builds and shares add-ons called “packages.”

Think of R as a brand-new, top-of-the-line smartphone. Itโ€™s powerful on its own, but it truly shines when you install apps. In R, these “apps” are called packages. For our forecasting mission, the most essential package is ‘forecast’. Developed by one of the world’s leading forecasting experts, Rob Hyndman, this package simplifies and automates many of the most complex parts of time series analysis. It contains the functions that will do the heavy lifting for us, making an advanced technique like ARIMA feel much less intimidating.

We will also use the ‘tseries’ package, which contains a crucial test for one of our first steps. You can easily install them in R by running install.packages("forecast") and install.packages("tseries").

The game plan: An ARIMA modelling workflow

We can’t just feed raw data into a machine and ask for a future number. We need a process. The one we’ll follow is the celebrated Box-Jenkins method, named after the statisticians who developed it. Itโ€™s an iterative, three-stage process that ensures we build a model that is both accurate and reliable. The three stages are: Identification, Estimation, and Diagnostic Checking.

Step 1: Identification (What kind of model fits our data?)

This is the most crucial (and most human) part of the process. We start by acting like a detective, looking for clues in our data.

Look at the data
Before you do anything else, plot your data. A simple line chart (using the plot() function in R) is the single most important diagnostic tool. We are looking for two main things:

  • Trend: Is the data generally moving upwards or downwards over time?
  • Seasonality: Are there repeating, predictable cycles (e.g., sales are always high in December, low in January)?

The concept of stationarity
ARIMA models have a very important requirement: they need the data to be stationary. A stationary time series is one whose statistical properties (like the mean and variance) are constant over time. Think of it this way: it’s incredibly difficult to forecast a river that is constantly getting wider and flowing faster. It’s much easier to predict a calm, steady canal. Our job is to turn our volatile river into a calm canal.

Most real-world economic data, like an exchange rate, is non-stationary. The INR-USD rate, for example, has a general upward trend over the last few decades. To fix this, we use a technique called differencing. This is the ‘I’ (Integrated) in ARIMA. Instead of modelling the price, we model the *change* in price from one day to the next. This new series (the day-to-day change) will likely be stationary, hovering around a mean of zero. The number of times we have to difference the data to make it stationary is our ‘d’ parameter.

While looking at a plot is a good start, we can formally test for stationarity using the Augmented Dickey-Fuller (ADF) test (from the ‘tseries’ package). In this test, a small p-value (e.g., < 0.05) gives us the confidence to say our data is stationary.

Finding ‘p’ and ‘q’ with ACF/PACF
Once our data is stationary (our ‘d’ is found), we need to find two more parameters. This is where ARIMA (AutoRegressive Integrated Moving Average) gets its other letters.

  • AR (AutoRegressive) – ‘p’: This part of the model assumes that the current value is dependent on its own past values. The ‘p’ is the number of past (lagged) values it depends on.
  • MA (Moving Average) – ‘q’: This part assumes that the current value is dependent on the “shocks” or forecast errors from the past. The ‘q’ is the number of past errors it depends on.

To find `p` and `q`, we look at two more plots:

  1. Autocorrelation Function (ACF): This plot shows the correlation of the series with its past values at different lags.
  2. Partial Autocorrelation Function (PACF): This plot is similar, but it “controls for” the correlations at shorter lags.

Reading these plots is a bit of an art, but the general rules of thumb are:

  • An AR(p) model (purely autoregressive) will have an ACF that “trails off” (slowly dies down) and a PACF that “cuts off” (drops to zero) after lag `p`.
  • An MA(q) model (purely moving average) will have an ACF that “cuts off” after lag `q` and a PACF that “trails off.”
  • An ARMA(p,q) model (a mix) will often have both plots trail off.

This step can be tricky and subjective. But don’t worry, R has a powerful shortcut for us, which we’ll see in the example.

Step 2: Estimation (Building the model)

Once we have our three parameters (p, d, q), this step is the easiest. We simply pass them to R’s arima() function. For example, if we identified an ARIMA(2,1,2) model, we’d just run: fit <- arima(my_data, order = c(2, 1, 2)) R will then use maximum likelihood estimation (MLE) to find the best possible coefficients for our AR(2) and MA(2) terms. It's like telling the master craftsman, "I need a cabinet with two shelves and two drawers," and R goes and builds it perfectly.

Step 3: Diagnostic checking (Did we build a good model?)

We're not done yet. After building the model, we must check our work. How? By looking at the residuals. The residuals are the "leftovers"-the difference between our model's predictions and the actual historical data. A good model should capture all the patterns in the data, leaving behind only random, unpredictable noise. This iterative cycle of identification, estimation, and checking is the core of the Box-Jenkins methodology.

The residuals should look like white noise. Think of it as the static on an old TV-no patterns, no trends, just pure randomness. We test this using a formal test called the Ljung-Box test. In R, the command is Box.test(fit$residuals, type = "Ljung-Box"). For this test, we want a high p-value (e.g., > 0.05). A high p-value means we *cannot* reject the null hypothesis that the residuals are independent and uncorrelated. In simple terms: a high p-value is good news, it means our residuals are just random noise, and our model is a good fit.

Let's make it real: Modelling the INR-USD exchange rate

Let's walk through this process with our real-world example. We want to forecast the daily INR-USD exchange rate. This is a classic problem in economics, and ARIMA is a well-suited tool for it.

Step 1 (Example): Identification

First, we load our data (let's say it's in a time series object called inr_usd_ts) and plot it. plot(inr_usd_ts) We'd immediately see a non-stationary upward trend. So, we run an ADF test: adf.test(inr_usd_ts) The p-value would be high (e.g., 0.7), confirming non-stationarity. So, we take the first difference: diff_data <- diff(inr_usd_ts) We plot diff_data and find it now hovers around zero. The ADF test on diff_data gives a tiny p-value (e.g., 0.01). Perfect. Our data is now stationary, and we've found our first parameter: d = 1.

Now, we look at the ACF and PACF plots for diff_data. acf(diff_data) pacf(diff_data)

Let's say we see both plots trailing off, with significant spikes at lags 1 and 2 in the PACF and lags 1 and 2 in the ACF. This suggests a mixed ARMA model. Based on our prompt, we might hypothesize an ARIMA(2,1,2) model.

The modern shortcut: auto.arima()

Guessing p and q from plots can be difficult. This is where the 'forecast' package becomes our best friend. We can use one powerful function to do this entire identification step for us: auto.arima(). auto_fit <- auto.arima(inr_usd_ts)

This single command runs a sophisticated algorithm that tests different combinations of p, d, and q. It automatically performs unit root tests to find `d`, and then minimizes an information criterion (like AICc) to find the "best" and most parsimonious `p` and `q`. It's a fantastic and reliable way to get a great model to start with. For our example, let's say auto.arima() also suggests an ARIMA(2,1,2) model.

Steps 2 & 3 (Example): Fitting and checking

Now we have our model, we can use the Arima() function from the 'forecast' package (note the capital 'A', which is a more robust version of the base `arima()` function). fit <- Arima(inr_usd_ts, order = c(2, 1, 2))

Next, we check the residuals. The 'forecast' package gives us another great shortcut: checkresiduals(fit)

This function provides a plot of the residuals, their ACF (which should have no significant spikes), and automatically runs the Ljung-Box test. We look at the test output: Ljung-Box test data: Residuals from ARIMA(2,1,2) Q* = 20.5, df = 21, p-value = 0.49

The p-value is 0.49. This is much higher than 0.05, so we are thrilled! This confirms our residuals are white noise. Our model is good. If the p-value were low, we would need to go back, check our (p,d,q) orders, and try a different model.

Diagnostic checking and forecasting in R (The payoff)

With our model (fit) built and validated, we can finally do what we came here for: forecast the future. This is the easiest part of all. We use the forecast() function.

Let's say we want to forecast the exchange rate for the next 30 days. my_forecast <- forecast(fit, h = 30)

This my_forecast object contains everything we need: the point forecasts (the single best guess), and also the 80% and 95% confidence intervals. The best way to see this is to plot it. plot(my_forecast)

The plot will show the historical data, followed by our 30-day forecast. The forecast won't be a single line; it will be a "fan" or "cone" that gets wider as it goes further into the future. This is a crucial and honest feature. It visually represents our uncertainty. Forecasting one day ahead is much easier than forecasting 30 days ahead, and the widening confidence intervals reflect that.

This entire, powerful workflow, from plotting raw data to forecasting, is the complete Box-Jenkins cycle. By following these steps-Identify, Estimate, Check, and Forecast-we've moved from just looking at data to building a robust statistical model that can help us navigate the uncertainties of the future.

What do you think? How might sudden economic events, like an RBI policy announcement, affect the accuracy of a simple ARIMA forecast? Beyond exchange rates, what other financial or economic data would you be interested in modelling with this process?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://en.wikipedia.org/wiki/Box%E2%80%93Jenkins_method
  2. https://www.rdocumentation.org/packages/stats/versions/3.6.2/topics/Box.test
  3. https://www.ies.gov.in/pdfs/SeminarPaper-ArushiGupta.pdf
  4. https://otexts.com/fpp2/arima-r.html
  5. https://www.geeksforgeeks.org/r-language/time-series-analysis-using-arima-model-in-r-programming/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Advanced Econometric Methods

1 Discrete Dependent Variable Models

  1. Introduction
  2. Qualitative Choice Analysis
  3. The Regression Approach
  4. The Latent Regression Approach
  5. The Probit Model
  6. The Logit Model
  7. Estimation and Inference

2 Censored and Truncated Regression Models

  1. Characteristics of Qualitative Response Models
  2. Tobit Model
  3. Truncated Regression Model
  4. Sample Selection Model
  5. Models with Multiple Choices

3 Autoregressive (AR) Models

  1. Structure of AR Models
  2. Reasons for Inclusion of Lags in AR Models
  3. Use of Lag Operator in AR Models
  4. Inter-temporal Effect of Shocks in AR Models
  5. Relevance of AR Models to Economic Theory
  6. Yule-Walker Equations in AR Models
  7. Estimation of Parameters of AR Model
  8. Use of AR Models in Financial Economics

4 Distributed Lag Models

  1. Distributed Lag Models
  2. Koyck Model
  3. Autoregressive Models
  4. A More General Dynamic Model
  5. Jorgensonโ€™s Rational Lag Model
  6. Partial Adjustment Model
  7. Adaptive Expectations Model
  8. Interpretation of Coefficients
  9. Estimation and Inference

5 Estimation of System of Equations

  1. Seemingly Unrelated Regression Equations (SURE)
  2. Generalized Least Squares (GLS)
  3. Feasible Generalized Least Squares (FGLS)
  4. Maximum Likelihood Estimates
  5. Hypothesis Testing
  6. Treating Autocorrelation
  7. Interrelated Factor Demand

6 Introduction to Simultaneous Equations Models

  1. Simultaneous Equations Model (SEM)
  2. Structural Form and Reduced Form
  3. Identification Problem
  4. Order Condition
  5. Rank Condition
  6. General Structure of SEM
  7. Simultaneity Bias

7 Estimation of Simultaneous Equations Models

  1. Limited Information Systems
  2. Full Information Systems

8 Specification Issues of Time Series Data Models

  1. Stochastic Process
  2. Detection of Unit Root โ€“ Graphical Examination
  3. Detection of Unit Root โ€“ Statistical Tests
  4. The KPSS Test
  5. Test for Unit Root in the Presence of Structural Break
  6. Relations among Non-Stationary Series
  7. Limitations of Engle-Granger Test

9 Modelling Univariate Time Series

  1. Autoregressive Models
  2. Moving Average Models
  3. ARMA Models
  4. Integrated Processes and the ARIMA Models
  5. Box-Jenkins Methodology
  6. ARIMA Modelling in Software R

10 Vector Auto-Regression (VAR) Models

  1. Specification and Estimation of VAR
  2. Uses of VAR
  3. Innovation Accounting
  4. Vector Autoregression of Non-Stationary Data

11 Modelling Volatility

  1. The Autoregressive Conditional Heteroscedasticity (ARCH) Model
  2. Properties of the ARCH Model
  3. Test for ARCH Effects
  4. Generalized-ARCH (GARCH) Model
  5. Extensions of the GARCH Model

12 Introduction to Panel Data Models

  1. Introduction
  2. Panel Data Models
  3. Fixed Effects Model
  4. Random Effects Model
  5. Choice between Fixed Effects and Random Effects Models
  6. Hausman Test

13 Dynamic Panel Data Analysis

  1. Static Panel Data Model
  2. Specification of Dynamic Panel Data Model
  3. Estimation Methods of Dynamic Panel data Models
  4. Arellano-Bond Estimator
  5. System-GMM Method of Estimation
  6. Problems with the Arellano-Bond Approach
  7. Maximum Likelihood Estimator

14 Introduction to Generalised Method of Moments Estimation

  1. Need for Generalized Method of Moments
  2. Additional Moments Restrictions and Generalized Method of Moments
  3. Leading Example of GMM: IV Regression in Overidentified Models
  4. Variance Estimation and Optimal GMM
  5. Estimating Optimal GMM โ€“ Two-Step GMM Estimator
  6. Test of Overidentifying Restrictions