Why deep learning is popular
Deep learning became popular because two things arrived together: huge amounts of data to learn from, and GPUs fast and cheap enough to train networks with many parameters on it.
Last updated: 05 Oct, 2026 · NumPy
Neural networks have been studied since 1958, so the interview question is why they took over only in the 2010s. The video gives two reasons, data and hardware; its notes add two more.
Following the growth of data
Go back to 2005. Orkut was the social network then, and Facebook followed. Facebook is built on web 2.0: you log in, store your data and interact with other people. Instagram, WhatsApp, LinkedIn and Twitter came next, and in all of them people post and interact. So data started to be generated exponentially.
Around 2008, storing that data efficiently became the problem, and big data engineering became a sought-after skill. By 2013 companies held huge amounts of data. Will they store it and keep it? No. They want to use it to give their customers a smoother experience and make them better, and that is only possible by learning from the data they already have. So AI became popular. Netflix is the example: it recommends movies from your profile and what you watch.
The years on the board mark eras, not launch dates: LinkedIn started in 2003, Orkut and Facebook in 2004, Twitter in 2006, WhatsApp in 2009 and Instagram in 2010.
Turning product data into revenue and adding GPUs
The video's own work at Panasonic shows the loop. Panasonic's ACs, TVs and refrigerators were already generating data. People use an AC badly, jumping between high and low temperatures, so a model built on the outside temperature and a usage profile can tell them how to run it and cut the electricity bill. That gives the customer a better experience, the model can be sold on a subscription basis, and the company generates revenue and makes better decisions.
The second reason is hardware advancement. NVIDIA makes GPUs, graphics processing units. A multi-layered neural network has many parameters and trains over many epochs, passes through the data, which takes a long time; GPUs train it fast, and their cost keeps falling.

Adding the reasons from the notes
The notes for this topic keep the two reasons and add two, with a chart that explains the first:

- Deep learning keeps improving with more data. A traditional algorithm levels off after a point; a large network keeps getting better as the data grows, so the data explosion favours it.
- It works in many domains: medical, e-commerce, retail, marketing.
- The frameworks are open source. TensorFlow comes from Google and PyTorch from Facebook (now Meta). Free tools grew a large community, and the community produced more research.
Counting the parameters a network trains
Why a GPU matters becomes clear once you count. Every neuron in a fully connected layer has one weight per input plus a bias, and every one of those numbers is updated on every training step.
Parameters in one layer
def dense_params(n_in, n_out):
# one weight per input for each neuron, plus one bias per neuron
return n_in * n_out + n_outComparing a small and a large network
networks = {
"3-5-4-3 (the board's multi-layer network)": [3, 5, 4, 3],
"784-512-512-10 (an image classifier)": [784, 512, 512, 10],
}
for name, sizes in networks.items():
total = sum(dense_params(a, b) for a, b in zip(sizes, sizes[1:]))
print(f"{name}: {total:,} parameters")
updates = 669_706 * (60_000 // 32) * 10 # parameters x batches of 32 x 10 epochs
print(f"weight updates over 10 epochs of 60,000 images: {updates:,}")3-5-4-3 (the board's multi-layer network): 59 parameters 784-512-512-10 (an image classifier): 669,706 parameters weight updates over 10 epochs of 60,000 images: 12,556,987,500
Reading the parameter counts
- 59 parameters for the board's 3-5-4-3 network: 20 + 24 + 15, small enough to train by hand.
- 669,706 parameters for a network that reads 28 × 28 pixel images, from two hidden layers of 512 neurons.
- Over 12 billion weight updates for ten epochs in batches of 32, and each update needs a forward and a backward pass. Those are matrix multiplications, the work GPUs do in parallel.
Classical machine learning vs deep learning
| Classical machine learning | Deep learning | |
|---|---|---|
| Data it needs | Works with hundreds or thousands of rows | Improves with much more data |
| Features | Made by hand from the raw data | Learned by the layers |
| Hardware | A CPU is enough | GPUs for large networks |
| Typical input | Tables of numbers | Images, text, audio, video |
| Explaining a prediction | Often easy (coefficients, tree splits) | Hard; a black box |
Where you use deep learning
- Products that already collect data, like the Panasonic ACs: recommendations, usage profiles, forecasts from sensor data.
- Images and video: medical scans, defect checks on a production line, the CNN part of this course.
- Text and speech: chatbots and translation, built on the RNN and transformer families.
Related
- Previous: AI vs ML vs DL vs data science
- Next: Installing TensorFlow
- Add a network
[11, 11, 7, 6, 1], the shape of the churn ANN later in the course, tonetworksand count its parameters. - Double the hidden layers of the image classifier to 1024 neurons each: by how much does the total grow?
- Change the batch size in
updatesfrom 32 to 256 (60_000 // 256) and see the total fall: 234 batches per epoch instead of 1,875.
You understood something today that you didn't yesterday.