Deep LearningTensorFlow 2.21 / Keras 3 · NumPy · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
41 small wins to finish your pathNext lesson →

Hidden layers and neurons with Keras Tuner

Keras Tuner is a hyperparameter search library that builds, trains and scores many versions of a Keras model to choose settings such as the number of hidden layers, the neurons in each layer and the learning rate.

Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras

In Building an ANN in Keras the video picks 11, 7 and 6 neurons by feel and says there are techniques to find how many hidden neurons a network needs. The Keras-Tuner notebook of the materials is that technique: it searches the number of layers, the units per layer and the learning rate on a regression problem, the air quality index.

Loading the air quality data

Real_Combine.csv holds daily weather readings, T (average temperature), TM (maximum), Tm (minimum), SLP (sea-level pressure), H (humidity), VV (visibility), V (wind speed) and VM (maximum wind speed), and the target PM 2.5, the concentration of fine particles in the air. Predicting a number makes it a regression, so the network ends in one linear neuron and learns with mean absolute error. The notebook takes every column but the last as X and the last as y, then splits 70/30. The run below reads the file from the Keras-Tuner repository and checks it for missing values first.

ExampleFrom the Keras-Tuner notebook, run on pandas 3 and scikit-learn 1.9.1
import math
import pandas as pd
from sklearn.model_selection import train_test_split

df=pd.read_csv("https://raw.githubusercontent.com/krishnaik06/Keras-Tuner/main/Real_Combine.csv")
print(df.shape)
print(df.head(3))
print("missing values:", df.isna().sum()[df.isna().sum() > 0].to_dict())

X=df.iloc[:,:-1] ## independent features
y=df.iloc[:,-1] ## dependent features
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)
print("NaN targets in the training set:", y_train.isna().sum())

df = df.dropna()                                   # the fix: drop the one bad row
X, y = df.iloc[:, :-1], df.iloc[:, -1]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)
print(X_train.shape, X_test.shape, "steps per epoch:", math.ceil(len(X_train) / 32))
  • 1,093 days and 9 columns: 8 features and PM 2.5.
  • One PM 2.5 value is missing, and it lands in the training set. The notebook never drops it, which matters when reading its training log below.
  • After dropna(), 764 training rows give 24 steps per epoch at Keras's default batch size of 32, the 24/24 of the notebook's log.

Defining the search space in build_model

The notebook imports RandomSearch from the tuner package:

python
import pandas as pd
from tensorflow import keras
from tensorflow.keras import layers
from kerastuner.tuners import RandomSearch

The package was renamed: kerastuner no longer exists, and current releases install as keras-tuner and import as keras_tuner.

A tuner needs a function that builds one model from a set of choices. build_model(hp) receives an hp object and asks it for every value it wants searched:

  • hp.Int('num_layers', 2, 20): the number of hidden layers, from 2 to 20.
  • hp.Int('units_' + str(i), 32, 512, step=32): the neurons of hidden layer i, one of 32, 64, ..., 512, chosen separately for every layer.
  • hp.Choice('learning_rate', [1e-2, 1e-3, 1e-4]): Adam's learning rate.
python
def build_model(hp):
    model = keras.Sequential()
    for i in range(hp.Int('num_layers', 2, 20)):
        model.add(layers.Dense(units=hp.Int('units_' + str(i),
                                            min_value=32,
                                            max_value=512,
                                            step=32),
                               activation='relu'))
    model.add(layers.Dense(1, activation='linear'))
    model.compile(
        optimizer=keras.optimizers.Adam(
            hp.Choice('learning_rate', [1e-2, 1e-3, 1e-4])),
        loss='mean_absolute_error',
        metrics=['mean_absolute_error'])
    return model

Creating a RandomSearch tuner

RandomSearch picks random combinations from that space. max_trials=5 tries five combinations, executions_per_trial=3 trains each one three times and averages the scores (the starting weights are random, so one run can be lucky), and objective='val_mean_absolute_error' ranks trials by their validation MAE, lower being better. Results are saved under directory/project_name.

python
tuner = RandomSearch(
    build_model,
    objective='val_mean_absolute_error',
    max_trials=5,
    executions_per_trial=3,
    directory='project',
    project_name='Air Quality Index')
ExampleOutput from the Keras-Tuner notebook (TensorFlow 2, keras-tuner as kerastuner, local run)
tuner.search_space_summary()

The summary lists units_0 and units_1 only because the space grows as models are built: a model with 10 layers registers units_0 to units_9 on its first call.

A search space of 2 to 20 layers, 32 to 512 units per layer in steps of 32 and three learning rates feeds five random trials; each trial builds a model and trains it three times, and the trial with the lowest average validation mean absolute error is the best.

tuner.search takes the same arguments as fit(), here 5 epochs with the test set as validation data, and runs every execution of every trial. The first execution of the first trial:

ExampleOutput from the Keras-Tuner notebook (TensorFlow 2, keras-tuner as kerastuner, local run)
tuner.search(X_train, y_train,
             epochs=5,
             validation_data=(X_test, y_test))

The training loss is nan at the end of every epoch while the validation loss is a number. That is the missing PM 2.5 value: the batch holding the NaN target has a NaN loss, and the running epoch average stays NaN from that batch on, which is why each line starts with a real loss and then switches to nan. The validation set has no missing value, so its MAE is still printed and keeps changing, but the training loss the log reports is useless. Dropping the row before the split fixes it.

Reading the best trial

ExampleOutput from the Keras-Tuner notebook (TensorFlow 2, keras-tuner as kerastuner, local run)
tuner.results_summary()

The best trial reached a validation MAE of 47.77 with a learning rate of 0.01 and 10 hidden layers of 192, 256 and then eight of 32 units; the next best scored 58.95. A trial with 3 layers still lists units_3 to units_14: those values were sampled for earlier, deeper models and are inactive in this one. The NumPy run below samples five trials the same way RandomSearch does and counts the parameters of each model, to show how different the candidates are.

ExampleRun with NumPy: random samples from the notebook's search space
import numpy as np

rng = np.random.default_rng(1)
def n_params(units, n_in=8):                       # 8 weather features in, 1 output
    sizes = [n_in, *units, 1]
    return sum(i * o + o for i, o in zip(sizes, sizes[1:]))

for trial in range(1, 6):                          # what RandomSearch samples
    n = rng.integers(2, 21)                        # hp.Int('num_layers', 2, 20)
    units = (rng.integers(1, 17, size=n) * 32).tolist()   # 32 to 512, step 32
    lr = rng.choice([1e-2, 1e-3, 1e-4])
    print(f"trial {trial}: {n:2d} layers, lr {lr}, {n_params(units):,} parameters")

best = [192, 256] + [32] * 8                       # the notebook's best trial
print("best trial: 10 layers", best, f"{n_params(best):,} parameters")
  • Five random models range from about 140,000 to 1.3 million parameters, from 6 to 20 layers: random search covers very different networks in a handful of trials.
  • The notebook's best model has 66,785 parameters, small next to most random picks, a sign that bigger is not automatically better here.
  • Five trials out of a space this large is a quick look, not an exhaustive search; more trials, or the Hyperband and BayesianOptimization tuners, explore it better.

Running the search on current Keras

With the current package name and the missing row dropped, the same search reads as follows. get_best_models returns the best model already trained, ready for evaluate() or predict().

python
import keras_tuner as kt                 # the package is keras-tuner, imported as keras_tuner

df = df.dropna()                         # drop the row with a missing PM 2.5
tuner = kt.RandomSearch(build_model, objective='val_mean_absolute_error',
                        max_trials=5, executions_per_trial=3,
                        directory='project', project_name='Air Quality Index')
tuner.search(X_train, y_train, epochs=5, validation_data=(X_test, y_test))
best_model = tuner.get_best_models(num_models=1)[0]

RandomSearch vs GridSearchCV

Keras Tuner RandomSearchGridSearchCV
Librarykeras_tuner, for Keras modelsscikit-learn, for scikit-learn models
How it picksrandom combinations, max_trials of themevery combination of the grid
Scoringa validation set, executions_per_trial runs averagedk-fold cross-validation
Search spacecan depend on itself (units per layer)a fixed grid
Covered inthis lessonHyperparameter tuning with GridSearchCV

Where you use Keras Tuner

  • Choosing the layer sizes of a tabular ANN like the churn network instead of picking 11, 7 and 6 by hand.
  • Tuning the learning rate and dropout rate together, with hp.Choice and hp.Float.
  • Searching CNN filters and kernel sizes with the same build_model pattern.
Watch out. A tuner ranks whatever scores it gets. In the notebook a single missing target turned every training loss into nan and the search still reported a best trial. Check df.isna().sum() and drop or fill missing values before searching.
Try it yourself
  • Change rng.integers(2, 21) to rng.integers(2, 4) in the sampling run and compare the parameter counts.
  • Remove the df = df.dropna() line in the data run and check which split the NaN target ends up in.
  • Count the parameters of the second-best trial (3 layers of 480, 480 and 512 units) with n_params.

Little by little, you're building something great.