Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Why deep learning is popular

Deep learning became popular because two things arrived together: huge amounts of data to learn from, and GPUs fast and cheap enough to train networks with many parameters on it.

Last updated: 05 Oct, 2026 · NumPy

Neural networks have been studied since 1958, so the interview question is why they took over only in the 2010s. The video gives two reasons, data and hardware; its notes add two more.

Following the growth of data

Data and the rise of AI · from the Deep Learning In-depth Tutorials in 5 Hours video · 14:31 to 19:00

Go back to 2005. Orkut was the social network then, and Facebook followed. Facebook is built on web 2.0: you log in, store your data and interact with other people. Instagram, WhatsApp, LinkedIn and Twitter came next, and in all of them people post and interact. So data started to be generated exponentially.

Around 2008, storing that data efficiently became the problem, and big data engineering became a sought-after skill. By 2013 companies held huge amounts of data. Will they store it and keep it? No. They want to use it to give their customers a smoother experience and make them better, and that is only possible by learning from the data they already have. So AI became popular. Netflix is the example: it recommends movies from your profile and what you watch.

The years on the board mark eras, not launch dates: LinkedIn started in 2003, Orkut and Facebook in 2004, Twitter in 2006, WhatsApp in 2009 and Instagram in 2010.

Turning product data into revenue and adding GPUs

The Panasonic example and GPUs · from the Deep Learning In-depth Tutorials in 5 Hours video · 19:00 to 22:02

The video's own work at Panasonic shows the loop. Panasonic's ACs, TVs and refrigerators were already generating data. People use an AC badly, jumping between high and low temperatures, so a model built on the outside temperature and a usage profile can tell them how to run it and cut the electricity bill. That gives the customer a better experience, the model can be sold on a subscription basis, and the company generates revenue and makes better decisions.

The second reason is hardware advancement. NVIDIA makes GPUs, graphics processing units. A multi-layered neural network has many parameters and trains over many epochs, passes through the data, which takes a long time; GPUs train it fast, and their cost keeps falling.

A timeline of why deep learning became popular: from 2005 social media made data grow exponentially, from 2008 big data stored it, and from 2013 companies used it in products so AI became popular; in a second lane NVIDIA GPUs train many-parameter networks fast while their cost falls; below, the Panasonic example runs from product data to a model, a subscription and revenue.

Adding the reasons from the notes

The notes for this topic keep the two reasons and add two, with a chart that explains the first:

Performance against amount of data: the deep learning curve keeps rising as data grows, while the traditional machine learning curve flattens early.
  • Deep learning keeps improving with more data. A traditional algorithm levels off after a point; a large network keeps getting better as the data grows, so the data explosion favours it.
  • It works in many domains: medical, e-commerce, retail, marketing.
  • The frameworks are open source. TensorFlow comes from Google and PyTorch from Facebook (now Meta). Free tools grew a large community, and the community produced more research.

Counting the parameters a network trains

Why a GPU matters becomes clear once you count. Every neuron in a fully connected layer has one weight per input plus a bias, and every one of those numbers is updated on every training step.

Parameters in one layer

python
def dense_params(n_in, n_out):
    # one weight per input for each neuron, plus one bias per neuron
    return n_in * n_out + n_out

Comparing a small and a large network

ExampleRun with Python 3.12
networks = {
    "3-5-4-3 (the board's multi-layer network)": [3, 5, 4, 3],
    "784-512-512-10 (an image classifier)": [784, 512, 512, 10],
}
for name, sizes in networks.items():
    total = sum(dense_params(a, b) for a, b in zip(sizes, sizes[1:]))
    print(f"{name}: {total:,} parameters")

updates = 669_706 * (60_000 // 32) * 10   # parameters x batches of 32 x 10 epochs
print(f"weight updates over 10 epochs of 60,000 images: {updates:,}")

Reading the parameter counts

  • 59 parameters for the board's 3-5-4-3 network: 20 + 24 + 15, small enough to train by hand.
  • 669,706 parameters for a network that reads 28 × 28 pixel images, from two hidden layers of 512 neurons.
  • Over 12 billion weight updates for ten epochs in batches of 32, and each update needs a forward and a backward pass. Those are matrix multiplications, the work GPUs do in parallel.

Classical machine learning vs deep learning

Classical machine learningDeep learning
Data it needsWorks with hundreds or thousands of rowsImproves with much more data
FeaturesMade by hand from the raw dataLearned by the layers
HardwareA CPU is enoughGPUs for large networks
Typical inputTables of numbersImages, text, audio, video
Explaining a predictionOften easy (coefficients, tree splits)Hard; a black box

Where you use deep learning

  • Products that already collect data, like the Panasonic ACs: recommendations, usage profiles, forecasts from sensor data.
  • Images and video: medical scans, defect checks on a production line, the CNN part of this course.
  • Text and speech: chatbots and translation, built on the RNN and transformer families.
Watch out. More data and a GPU do not make deep learning the right choice for every problem. On a small table the curves cross the other way: a classical model trains in a second, is easier to explain and often scores as well. Try a simple model first.
Try it yourself
  • Add a network [11, 11, 7, 6, 1], the shape of the churn ANN later in the course, to networks and count its parameters.
  • Double the hidden layers of the image classifier to 1024 neurons each: by how much does the total grow?
  • Change the batch size in updates from 32 to 256 (60_000 // 256) and see the total fall: 234 batches per epoch instead of 1,875.

You understood something today that you didn't yesterday.