When people hear “data science”, many picture complex formulas, artificial intelligence and something “for mathematicians only”. In reality the field is learnt step by step, and each step is a useful skill in its own right: someone who can work with spreadsheets can put together a report, someone who knows SQL can pull the right data from a database, and someone who knows Python can automate repetitive work. In this article we look at how the roles of data analyst, data scientist and ML engineer differ, how much maths you really need, the sensible order for learning everything from Excel and statistics to machine learning, how to use open datasets and what kind of projects to build for a portfolio. At the end you will find a sample month-by-month plan, common mistakes, a practice exercise and a checklist.
Data analyst, data scientist and ML engineer: what is the difference
All three roles work with data, but they ask different questions and produce different results. Job adverts sometimes blur these titles, so look at the list of responsibilities rather than the title itself.
- Data analyst answers the question “what happened and why?”. They work with spreadsheets, SQL and visualisation tools, prepare reports and dashboards and give management clear conclusions. For example, they find out which branch of a retail chain saw sales fall, and why.
- Data scientist also tackles the question “what is likely to happen?”. They apply statistical methods and machine learning models: forecasting demand, grouping customers, assessing the results of experiments.
- ML engineer turns a finished model into a working system. The role is closer to that of a programmer: code quality, servers, and the speed and reliability of the model.
A simple comparison: Dilshod works as an analyst for a chain of shops and prepares a sales report every week. His colleague Zarina builds a model that forecasts which products will be in higher demand next month. A third colleague, Bekzod, connects that model to the warehouse software so that it runs by itself every day.
For beginners, the most realistic way in is the data analyst route. It needs less maths, the results show quickly, and later it gives you a solid foundation for moving into data science. If you would like to understand what data analysis is through simple examples, start with our article on data analysis.
How much maths you need: a realistic view
Many people are put off this field because “I am not good at maths”. The realistic answer is that school maths and basic statistics are enough to start, while deeper maths can be learnt later, as and when you need it.
For a data analyst:
- working comfortably with percentages, ratios and averages;
- understanding the difference between the mean, median, mode, minimum and maximum;
- knowing what spread (standard deviation) tells you;
- realising that correlation and cause and effect are not the same thing.
For a data scientist, in addition:
- the basics of probability theory;
- elements of linear algebra: vectors, matrices and the idea of multiplying them;
- the concepts of a derivative and a gradient, so you understand how a model “learns”.
The key point: libraries do the calculations for you, so what you need is not to solve formulas by hand but to interpret results correctly. For example, noticing that a model’s high “accuracy” may be misleading because the data is unbalanced is exactly what mathematical thinking means here. If you want to start statistics from scratch, our article on statistics basics will help.
Stage one: Excel, statistics and SQL
Excel and spreadsheets
A spreadsheet is the most convenient “laboratory” for working with data. You can see the data with your own eyes and spot mistakes straight away. At this stage, learn the following:
- Building a clean table: one value per cell, one type of data per column.
- Formulas: SUM, AVERAGE, COUNTIF, IF, XLOOKUP (or VLOOKUP).
- Sorting and filtering: Data → Sort, Data → Filter.
- Pivot tables: Insert → PivotTable, for grouping, counting and totalling.
- Charts: Insert → Chart.
These skills are the foundation for every later stage.
Statistics basics
Learn statistics alongside spreadsheets. For example, enter the monthly electricity use of families in one mahalla in kilowatt-hours (illustrative, made-up figures) and compare the mean with the median. You will see how strongly a single very large value “pulls” the mean towards it, which is one of the most important lessons in statistics.
SQL
In organisations, data is usually stored not in spreadsheets but in databases, and it is retrieved with SQL queries. Most data analyst job adverts ask for SQL. The order to learn it in:
- SELECT, FROM, WHERE: choosing the columns and rows you need.
- ORDER BY and LIMIT: sorting and limiting results.
- GROUP BY and aggregate functions (COUNT, SUM, AVG).
- JOIN: linking several tables.
- Subqueries and window functions: the next level.
A simple query:
SELECT region, COUNT(*) AS order_count
FROM orders
GROUP BY region
ORDER BY order_count DESC;
This query counts the orders for each region and lists them starting from the largest. To learn SQL from scratch, see our article on SQL basics.
Stage two: Python, pandas and visualisation
Once you are comfortable with spreadsheets and SQL, move on to Python. It automates repetitive work, copes with large volumes of data and opens the door to machine learning. If programming is completely new to you, first get the basics from our article on your first steps in Python: variables, lists, conditions, loops and functions.
A convenient environment for learning is Jupyter Notebook or Google Colab. Colab runs in the browser and needs nothing installed on your computer: click File → New notebook and start writing code.
pandas
pandas is the main library for working with tabular data. With it, much of what you do by hand in Excel takes just a few lines of code:
import pandas as pd
df = pd.read_csv("savdo.csv")
print(df.head())
print(df.describe())
print(df.groupby("viloyat")["miqdor"].sum())
Here read_csv reads the file, head shows the first rows, describe prints the basic statistics, and groupby groups the data by region and adds it up. The main operations to learn are selecting columns, filtering, finding and filling empty values, changing a column’s type and combining tables.
Visualisation
Data visualisation is the craft of presenting what you have found so that others can understand it. In Python the basic library is matplotlib, and seaborn is built on top of it and makes neater charts easier to draw. The main chart types are bar charts (comparing categories), line charts (change over time), histograms (distribution) and scatter plots (the relationship between two variables).
The rules of a good chart: a clear title, labelled axes, units of measurement and no unnecessary colours. One chart, one message. If you cannot state its conclusion in one sentence, simplify it.
Stage three: machine learning basics
Machine learning means teaching a computer to find patterns from examples without writing explicit rules for it. It is the most practical part of the field of artificial intelligence. For a beginner, two core ideas matter most.
Supervised learning. The model is given examples with a question and the correct answer. For example, data on pupils from previous years: attendance, hours spent on homework and final exam scores. The model learns the relationship between them and predicts the score for a new pupil. There are two types of task here: regression (predicting a number) and classification (assigning a category, for example whether an email is spam or not).
Unsupervised learning. There are no correct answers, and the model finds natural groups in the data by itself. For example, splitting an online shop’s customers into several groups according to their shopping habits (clustering).
The main library to start with in Python is scikit-learn. Almost every model in it is used in the same way:
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2)
model = LinearRegression()
model.fit(X_train, y_train)
print(model.score(X_test, y_test))
The key concepts are splitting data into training and test sets (the model must not see the test set while it learns), training the model (fit), making predictions (predict) and evaluating quality. The most dangerous mistake is a model that memorises the training data and then performs poorly on new data. This is called overfitting.
At this stage, stick to simple models: linear regression, logistic regression, decision trees and k-means clustering. Move on to neural networks and deep learning once your foundations are solid.
Projects, a portfolio and open datasets
Employers want to see what you can actually do rather than a list of certificates. So at the end of each stage, build a small project and collect them in a portfolio.
Where to find a dataset
A dataset is a collection of data prepared for analysis. There are open sources for practice:
- Kaggle: thousands of datasets uploaded by users around the world, along with write-ups and learning competitions. Reading other people’s work is a good way to learn, too.
- data.gov.uz: Uzbekistan’s open data portal, where data published by various public bodies is made available. A project based on local data looks particularly interesting in a portfolio.
- Your own data: household spending, a reading diary, the match results of your mahalla football team. Before publishing personal data, remove names, phone numbers and similar details.
What a good project looks like
- A clear question. Not “I analysed a dataset” but “In which months do library visits drop, and why?”
- Data cleaning. Describe how you dealt with empty values, duplicates and incorrect formats.
- Analysis and charts. Three to five meaningful charts are enough.
- A conclusion. In plain language: what you found, what the limitations are, and what could be done next.
- Publishing. Put the code on GitHub and briefly explain the project in a README file.
We have written in more detail about putting together a portfolio in our article on creating a portfolio.
Why English matters
Most of the documentation, library guides, question-and-answer forums and new courses in this field are in English. Even to look up an error message you need to be able to read and write in English. You do not need to speak fluently, but being able to read and understand technical text is a huge advantage. Even 15–20 minutes of reading technical material a day makes a noticeable difference.
A sample month-by-month learning plan
The table below is illustrative. It is designed for someone starting programming from scratch and studying for roughly 8–10 hours a week. If you can give it more time you will move faster; with less time, more slowly. What matters is not speed but consistency.
| Month | Topic | Outcome of the stage |
|---|---|---|
| 1 | Excel, clean tables, formulas, PivotTable | A personal spending analysis spreadsheet |
| 2 | Statistics basics, charts | A short report comparing the mean and median |
| 3 | SQL: SELECT, GROUP BY, JOIN | A set of 15–20 practice queries |
| 4 | Python basics | A small automation script |
| 5 | pandas, data cleaning | Cleaning and describing an open dataset |
| 6 | matplotlib, seaborn, visualisation | A complete analysis project (question, charts, conclusion) |
| 7–8 | scikit-learn, supervised and unsupervised learning | One prediction project and one clustering project |
| 9 onwards | Portfolio, GitHub, CV, revision | Three or four well-documented projects |
At the end of each month, ask yourself: “Could I explain what I learnt this month to someone else and apply it to a small task?” If the answer is no, do not rush on to the next topic.
Common mistakes
- Only watching courses. Videos give a feeling of understanding, but skills only form when you write code yourself. After every lesson, repeat what you learnt with your own example.
- Learning without projects. Finishing ten courses without doing a single independent project. Employers look at results, not certificates.
- Skipping the basics. Starting straight with neural networks. Without Excel, SQL and statistics it is hard to evaluate a model properly.
- Everything at once. Starting Python, R, five libraries and three courses in parallel. Choose one path and follow it through to the end.
- Not checking the data. Building a model on uncleaned data. Experienced specialists spend much of their time understanding and cleaning data.
- Not being able to explain the result. Saying “the model is highly accurate” is not enough. You need to be able to say in plain language what the result means for the organisation or for people.
- Expecting quick results. This field takes several months of regular work, or longer. No course can guarantee you a job; the outcome depends on your own practical work.
Practice exercise and checklist
Exercise: analysing a mahalla library
Imagine that Nilufar works at her mahalla library and wants to know which books are read most often. The data is illustrative, and you will create it yourself.
- Create 30–40 rows of data in a spreadsheet: date, book title, genre, reader’s age group, day of return.
- Use a PivotTable in Excel to count how many books were borrowed in each genre.
- For each genre, compare the mean and median number of days before a book is returned.
- Save the table in CSV format (File → Save As → CSV).
- In Google Colab, read the file with pandas and repeat steps 2–3 in code.
- Draw a bar chart by genre with matplotlib.
- Write a three-sentence conclusion: what you found, what the limitations are and what you would recommend to the library.
- Put everything on GitHub and write a README.
Checklist
- I can explain the difference between a data analyst, a data scientist and an ML engineer in my own words.
- I can calculate and interpret the mean, median and standard deviation.
- I can build a pivot table and a chart in Excel.
- I can write SQL queries with GROUP BY and JOIN.
- I can read, clean and group a CSV file with pandas.
- I can explain the difference between supervised and unsupervised learning with an example.
- At least one of my projects is on GitHub with a README.
- I can read and understand technical text in English.
Data science is a path built step by step, not in a single day. Each stage gives you a useful skill in its own right, so even in the first months you will start applying what you learn at work or in everyday life.
If you would like to follow this path systematically with a teacher, take a look at our association’s free Data analytics and Data Science programmes, as well as our other training programmes. When you are ready, submit an application and our specialists will get in touch.