Researching and Modeling Data: Statistics and Forecast Models

19 İyul 2026·👁 7 views

The third and fourth phases of the OSEMN framework — research (explore) and modeling — are where the true power of the dataset emerges. You recognize the dataset in research, and predict the future in modeling. In this article, we explain the «language» of the dataset, the main statistics and the main model types.

This post what is data analytics a continuation of its guide.

Learning the language of data

Think of data literacy as French or Spanish: the more you understand, the better you use it. Start by asking questions in the dataset:

  • How many data sources do I have, and how many files each?
  • What is the size of the files — a few kilobytes or gigabytes?
  • How many rows and columns are there in each dataset?
  • What type of data is there in each column?

Data types in columns can be: digital, text, boolean («true»/«false»), or calendar (date and time). An important allocation: is the data digital (numerical/quantitative) or categorical (categorical)? Categories data fall into different groups-for example, ice cream flavors or marketing channels (Email, TV, Social Media).

Research in practice: directive work

In the research phase, you're actually looking for patterns, trends, and anomalies below the surface of the dataset. Practically speaking, you start with a few questions: which columns are related to each other? Which variables have increased or decreased over time? What values go beyond expectations (outlier)? The answers to these questions often lead to important insights that can't be expected.

For example, a marketing channel column might have three groups — Email, TV, and Social Media. If you count how many times each value appears (frequency), let's say you find 3 Emails, 1 Social Media, and 4 TV records. This simple calculation shows which channel dominates the dataset. Research starts with such small discoveries and goes deeper.

Summary statistics: a brief picture of the dataset

Datasets are often so large that not every line can be viewed. That's why data analysts use summary statistics — mathematical tools that aggregate large datasets into a single digit. For category data, you calculate how many groups there are and the frequency (frequency) of each value. For digital data, you can:

  • Minimum and maximum values.
  • Median (midpoint) and mode (most recurring value).
  • Mean (average) and standard deviation (how far the values differ from the mean).

It's also helpful to look directly at the raw data: the first few lines («head»), the last few lines («tail»), or random sampling. Understanding the language of the dataset is not a theoretical exercise — it is a key tool that helps to produce insight in the analysis work.

Visualization, distribution, and connections

The research phase takes place not only in numbers, but also in expectation. Visualization allows you to quickly see the patterns inside the dataset — like a bar chart, line graph, histogram, and scatter diagram (scatter plot). A chart often speaks more than a long one.

There are two main areas of investigation. First — check the data distribution (s): how are the values of a variable spread? Are most of them in the middle or prone to flies? The histogram shows this. Secondly — checking the data relationship (relationships): are the two variables related to each other? For example, does sales increase as ad spend increases? A scattering diagram and correlation reveal this. This relationship is the basis of further modeling.

The research also has the concept of «feature engineering» (sign engineering). This is to create new, more useful variables from an existing dataset. For example, calculating an “age” from the “date of birth” column or subtracting a “day of the week”from the “date”. The right signs can significantly improve the model's accuracy.

What are models and why are they used?

Before building a model, let's start with a question: what is a model? Models are mathematical instruments that recognize hidden patterns in the dataset that may not be visible to a person. The most common use is to predict the future based on the past dataset. The model can also be simple (calculate the middle of a dataset and say the future will be average) and complex (machine learning models that predict the share price).

Yaxşı xəbər budur ki, modeli istifadə etmək üçün onu qurmağı bilmək şərt deyil. Analitiklər çox vaxt data alimləri komandası ilə işləyir. Ən çox istifadə olunan sadə modellərdən biri xətti reqressiyadır (linear regression) — bir dəyişənin dəyərini digəri ilə proqnozlaşdıran xətt. Məsələn, evin qiymətini onun sahəsinə (kvadrat fut) görə proqnozlaşdırmaq: 3000 kvadrat futluq ev üçün x oxunda həmin nöqtəni tapıb proqnoz xəttinə qədər qalxırsan. Sahə artdıqca qiymət də artır — məntiqlidir.

Predictions are never 100% accurate. The «shaded area» in the model shows a margin of error with 95% confidence — meaning we're 95% sure the actual value will be in that area. As the statistician George Box famously said: "All models are wrong, but some are useful.» While the model may not be perfect, it can be as accurate as creating business value with the right data.

How do models work?

Building a model is essentially giving the machine the ability to “learn” from a past dataset. In the home price example, you give the machine lots of home data (fields and prices). The machine looks at these points and finds the line that best fits between them — this line is the model. Then, for a new home (e.g. 3,000 square feet), use that line to make a price forecast. The dots in the graph are real examples of what the machine learns, and the line is its prediction.

There's a difference between building and using a model. Model building (training) is about giving a machine a dataset and teaching examples; using a model is about asking a new question and getting a prediction on the model being taught. As a data analyst, you often play a role not only in building a model, but in selecting and using it correctly. The important thing is to understand what the model means and how much to trust in its results.

Being able to quantify uncertainty is the main strength of the model. The 95% confidence margin in the home price model tells us how much deviation we can expect from the forecast. This allows us to make smart business decisions — because we know how much we can rely on forecasting.

Model types

To select a model, first ask: which question do I want answered?

  • Regression (regression) models answer the questions «how much?» — digitally. Sales or click-through rate forecast, share price.
  • Classification models predict group/class — categorical. Binary («true»/«false») or multi-individual (picture a dog, cat, or human?). Health disease estimates, whether or not a customer passes to a competitor in marketing.
  • Clustering algorithms divide data into similar segments. For example, allocating the customer base to niche audiences such as “sports fans ages 18-25”.

Three popular examples of algorithms are linear regression (simple connections, requires little data), decision trees (decision trees — the possibility of debt repayment based on binary decisions, e.g. credit scoring; when combined, it becomes a «random forest»), and neural networks (neural networks — complex connections, real estate, self-driving cars, art creations; require a lot of data and computing power). The key to success is choosing the model that fits the problem.

How to choose the right model

Hər model müəyyən problem növü üçün digərlərindən daha effektivdir, ona görə seçim vacibdir. Seçimə sualla başla: rəqəmsal cavab istəyirsən (reqressiya), qrup proqnozu (klassifikasiya), yoxsa oxşar seqmentlər tapmaq (klasterləmə)? Sonra datanın həcmini və mürəkkəbliyini nəzərə al: az data ilə sadə problem üçün xətti reqressiya idealdır; orta mürəkkəblik üçün qərar ağacları (az hesablama gücü tələb edir); çox mürəkkəb, çoxdəyişənli problem üçün isə neyron şəbəkələri. Hər modelin öz güclü və zəif tərəflərini anlamaq layihən üçün ən yaxşı seçimi etməyə kömək edir.

Remember, it's not about looking for the perfect model, it's about building a useful model. As George Box said, all models are wrong, but some are useful. A simple, intelligible model is often more valuable than a complex, “black box” model because you can explain its results and communicate them to decision makers.

Frequently asked questions (FAQs)

Do you need to be a developer to build a model?

Not a prerequisite. Many models are easy to apply; analysts often work with data scientists. The key skill is choosing the right model and using it correctly.

What is linear regression?

A line that predicts one variable with another. Ideal for simple problems and requires little data.

What is the difference between regression and classification?

Regression is a digital response («how much?»); classification is a categorical response (which group?).

When is a neural network in use?

When it comes to complex, multi-tenant connections — real estate, a self-driving car, art creation. It requires a lot of data and computing power.

What does «All models are wrong» mean?

The statistic is George Box's word: the model never perfectly reflects reality, but with the right data, it can be useful enough to create business value.

What are summary statistics?

Mathematical tools that aggregate a large dataset into a single digit — such as mean (average), median, mode, minimum, maximum, and standard deviation. Gives a quick overview of the dataset.

What is Feature Engineering?

Creating new, more useful variables from an existing dataset — for example, calculating the age from the date of birth. Correct signs improve model accuracy.

Where is the clustering used?

To divide the customer base into similar segments — for example, “sports fans aged 18-25” — so you can build more focused marketing for each group.

What is data distribution?

Shows how the values of a variable are spread out — most are in the middle or prone to flies. A histogram shows this visually.

What is a decision tree?

A model that divides the dataset based on binary decisions. For example, credit scoring predicts the likelihood of debt repayment due to income and employment status. When combined, it is a "random forest".

What is the difference between explore and model stages?

Explore datasets are about recognizing and seeing patterns; the model is about mathematically translating those patterns into future predictions. One describes, the other predicts.

Conclusion: research and modeling together

Research and modeling complement each other. First, you recognize the dataset, learn its language, and see the patterns and connections in it (explore). Then you turn these patterns into a mathematical model and predict the future (model). It's dangerous to build a model without good research — because you won't know which variables are important. Without a good model, research just describes the past and doesn't predict the future. Together, the two make the dataset a real decision tool.

An important reminder: models help us analyze the dataset first so we can interpret it and predict the future. Modeling is not a goal in and of itself — the goal is to think and decide. That's why the next step — interpreting and communicating results — is so important.

You don't need complicated models to get started. You can learn a lot with simple summary statistics and linear regression. As the experience evolves, you move to more complex models. The most important thing is to know the dataset well and choose the simplest model that fits the problem; complexity doesn't always mean better results.

Next step: learn how to interpret model results and communicate them to others.

👉 Next Post data storytelling

Tural Rəhimov

About the author

Tural Rəhimov

Digital Marketing Manager — UM Azerbaijan

Digital marketing manager at Universal McCann (UM Azerbaijan). Experienced in Google Ads, Meta Ads, TikTok Ads and media planning. I help brands grow online.

More about me →

Need digital marketing support for your business?

Let's talk — I can help with everything from strategy to ad management.

Message me