Hidden layers and neurons with Keras Tuner
Keras Tuner is a hyperparameter search library that builds, trains and scores many versions of a Keras model to choose settings such as the number of hidden layers, the neurons in each layer and the learning rate.
Last updated: 05 Oct, 2026 · TensorFlow 2 / Keras
In Building an ANN in Keras the video picks 11, 7 and 6 neurons by feel and says there are techniques to find how many hidden neurons a network needs. The Keras-Tuner notebook of the materials is that technique: it searches the number of layers, the units per layer and the learning rate on a regression problem, the air quality index.
Loading the air quality data
Real_Combine.csv holds daily weather readings, T (average temperature), TM (maximum), Tm (minimum), SLP (sea-level pressure), H (humidity), VV (visibility), V (wind speed) and VM (maximum wind speed), and the target PM 2.5, the concentration of fine particles in the air. Predicting a number makes it a regression, so the network ends in one linear neuron and learns with mean absolute error. The notebook takes every column but the last as X and the last as y, then splits 70/30. The run below reads the file from the Keras-Tuner repository and checks it for missing values first.
import math
import pandas as pd
from sklearn.model_selection import train_test_split
df=pd.read_csv("https://raw.githubusercontent.com/krishnaik06/Keras-Tuner/main/Real_Combine.csv")
print(df.shape)
print(df.head(3))
print("missing values:", df.isna().sum()[df.isna().sum() > 0].to_dict())
X=df.iloc[:,:-1] ## independent features
y=df.iloc[:,-1] ## dependent features
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)
print("NaN targets in the training set:", y_train.isna().sum())
df = df.dropna() # the fix: drop the one bad row
X, y = df.iloc[:, :-1], df.iloc[:, -1]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=0)
print(X_train.shape, X_test.shape, "steps per epoch:", math.ceil(len(X_train) / 32))(1093, 9)
T TM Tm SLP H VV V VM PM 2.5
0 7.4 9.8 4.8 1017.6 93.0 0.5 4.3 9.4 219.720833
1 7.8 12.7 4.4 1018.5 87.0 0.6 4.4 11.1 182.187500
2 6.7 13.4 2.4 1019.4 82.0 0.6 4.8 11.1 154.037500
missing values: {'PM 2.5': 1}
NaN targets in the training set: 1
(764, 8) (328, 8) steps per epoch: 24- 1,093 days and 9 columns: 8 features and PM 2.5.
- One PM 2.5 value is missing, and it lands in the training set. The notebook never drops it, which matters when reading its training log below.
- After
dropna(), 764 training rows give 24 steps per epoch at Keras's default batch size of 32, the 24/24 of the notebook's log.
Defining the search space in build_model
The notebook imports RandomSearch from the tuner package:
import pandas as pd
from tensorflow import keras
from tensorflow.keras import layers
from kerastuner.tuners import RandomSearchThe package was renamed: kerastuner no longer exists, and current releases install as keras-tuner and import as keras_tuner.
A tuner needs a function that builds one model from a set of choices. build_model(hp) receives an hp object and asks it for every value it wants searched:
- hp.Int('num_layers', 2, 20): the number of hidden layers, from 2 to 20.
- hp.Int('units_' + str(i), 32, 512, step=32): the neurons of hidden layer i, one of 32, 64, ..., 512, chosen separately for every layer.
- hp.Choice('learning_rate', [1e-2, 1e-3, 1e-4]): Adam's learning rate.
def build_model(hp):
model = keras.Sequential()
for i in range(hp.Int('num_layers', 2, 20)):
model.add(layers.Dense(units=hp.Int('units_' + str(i),
min_value=32,
max_value=512,
step=32),
activation='relu'))
model.add(layers.Dense(1, activation='linear'))
model.compile(
optimizer=keras.optimizers.Adam(
hp.Choice('learning_rate', [1e-2, 1e-3, 1e-4])),
loss='mean_absolute_error',
metrics=['mean_absolute_error'])
return modelCreating a RandomSearch tuner
RandomSearch picks random combinations from that space. max_trials=5 tries five combinations, executions_per_trial=3 trains each one three times and averages the scores (the starting weights are random, so one run can be lucky), and objective='val_mean_absolute_error' ranks trials by their validation MAE, lower being better. Results are saved under directory/project_name.
tuner = RandomSearch(
build_model,
objective='val_mean_absolute_error',
max_trials=5,
executions_per_trial=3,
directory='project',
project_name='Air Quality Index')tuner.search_space_summary()Search space summary |-Default search space size: 4 num_layers (Int) |-default: None |-max_value: 20 |-min_value: 2 |-sampling: None |-step: 1 units_0 (Int) |-default: None |-max_value: 512 |-min_value: 32 |-sampling: None |-step: 32 units_1 (Int) |-default: None |-max_value: 512 |-min_value: 32 |-sampling: None |-step: 32 learning_rate (Choice) |-default: 0.01 |-ordered: True |-values: [0.01, 0.001, 0.0001]
The summary lists units_0 and units_1 only because the space grows as models are built: a model with 10 layers registers units_0 to units_9 on its first call.

Running the search
tuner.search takes the same arguments as fit(), here 5 epochs with the test set as validation data, and runs every execution of every trial. The first execution of the first trial:
tuner.search(X_train, y_train,
epochs=5,
validation_data=(X_test, y_test))Epoch 1/5 ... 24/24 [==============================] - ETA: 0s - loss: 116.8132 - mean_absolute_error: 116.813 - ETA: 0s - loss: 76.3447 - mean_absolute_error: 76.3447 - 0s 18ms/step - loss: nan - mean_absolute_error: nan - val_loss: 63.4238 - val_mean_absolute_error: 63.4238 Epoch 2/5 24/24 [==============================] - ETA: 0s - loss: 72.3939 - mean_absolute_error: 72.393 - ETA: 0s - loss: nan - mean_absolute_error: nan - 0s 16ms/step - loss: nan - mean_absolute_error: nan - val_loss: 62.7431 - val_mean_absolute_error: 62.7431 Epoch 3/5 24/24 [==============================] - ETA: 0s - loss: 78.7923 - mean_absolute_error: 78.792 - ETA: 0s - loss: nan - mean_absolute_error: nan - 0s 19ms/step - loss: nan - mean_absolute_error: nan - val_loss: 47.5451 - val_mean_absolute_error: 47.5451 Epoch 4/5 24/24 [==============================] - ETA: 0s - loss: 48.1353 - mean_absolute_error: 48.135 - ETA: 0s - loss: nan - mean_absolute_error: nan - 0s 4ms/step - loss: nan - mean_absolute_error: nan - val_loss: 72.8017 - val_mean_absolute_error: 72.8017 Epoch 5/5 24/24 [==============================] - ETA: 0s - loss: 58.7824 - mean_absolute_error: 58.782 - ETA: 0s - loss: nan - mean_absolute_error: nan - 0s 4ms/step - loss: nan - mean_absolute_error: nan - val_loss: 55.2045 - val_mean_absolute_error: 55.2045 ...
The training loss is nan at the end of every epoch while the validation loss is a number. That is the missing PM 2.5 value: the batch holding the NaN target has a NaN loss, and the running epoch average stays NaN from that batch on, which is why each line starts with a real loss and then switches to nan. The validation set has no missing value, so its MAE is still printed and keeps changing, but the training loss the log reports is useless. Dropping the row before the split fixes it.
Reading the best trial
tuner.results_summary()Results summary |-Results in project\Air Quality Index |-Showing 10 best trials |-Objective(name='val_mean_absolute_error', direction='min') Trial summary |-Trial ID: bd9d67176a48257acd476a792a58666c |-Score: 47.77180608113607 |-Best step: 0 Hyperparameters: |-learning_rate: 0.01 |-num_layers: 10 |-units_0: 192 |-units_1: 256 |-units_2: 32 |-units_3: 32 |-units_4: 32 |-units_5: 32 |-units_6: 32 |-units_7: 32 |-units_8: 32 |-units_9: 32 Trial summary |-Trial ID: 94639d964caf10b4baa7059e8d927d02 |-Score: 58.95306905110677 |-Best step: 0 Hyperparameters: |-learning_rate: 0.001 |-num_layers: 3 |-units_0: 480 |-units_1: 480 |-units_10: 352 |-units_11: 416 |-units_12: 384 |-units_13: 320 |-units_14: 160 |-units_2: 512 |-units_3: 64 |-units_4: 448 |-units_5: 128 |-units_6: 416 |-units_7: 64 |-units_8: 352 |-units_9: 192 ...
The best trial reached a validation MAE of 47.77 with a learning rate of 0.01 and 10 hidden layers of 192, 256 and then eight of 32 units; the next best scored 58.95. A trial with 3 layers still lists units_3 to units_14: those values were sampled for earlier, deeper models and are inactive in this one. The NumPy run below samples five trials the same way RandomSearch does and counts the parameters of each model, to show how different the candidates are.
import numpy as np
rng = np.random.default_rng(1)
def n_params(units, n_in=8): # 8 weather features in, 1 output
sizes = [n_in, *units, 1]
return sum(i * o + o for i, o in zip(sizes, sizes[1:]))
for trial in range(1, 6): # what RandomSearch samples
n = rng.integers(2, 21) # hp.Int('num_layers', 2, 20)
units = (rng.integers(1, 17, size=n) * 32).tolist() # 32 to 512, step 32
lr = rng.choice([1e-2, 1e-3, 1e-4])
print(f"trial {trial}: {n:2d} layers, lr {lr}, {n_params(units):,} parameters")
best = [192, 256] + [32] * 8 # the notebook's best trial
print("best trial: 10 layers", best, f"{n_params(best):,} parameters")trial 1: 10 layers, lr 0.001, 788,129 parameters trial 2: 7 layers, lr 0.0001, 313,409 parameters trial 3: 16 layers, lr 0.001, 992,033 parameters trial 4: 6 layers, lr 0.01, 139,905 parameters trial 5: 20 layers, lr 0.001, 1,310,561 parameters best trial: 10 layers [192, 256, 32, 32, 32, 32, 32, 32, 32, 32] 66,785 parameters
- Five random models range from about 140,000 to 1.3 million parameters, from 6 to 20 layers: random search covers very different networks in a handful of trials.
- The notebook's best model has 66,785 parameters, small next to most random picks, a sign that bigger is not automatically better here.
- Five trials out of a space this large is a quick look, not an exhaustive search; more trials, or the Hyperband and BayesianOptimization tuners, explore it better.
Running the search on current Keras
With the current package name and the missing row dropped, the same search reads as follows. get_best_models returns the best model already trained, ready for evaluate() or predict().
import keras_tuner as kt # the package is keras-tuner, imported as keras_tuner
df = df.dropna() # drop the row with a missing PM 2.5
tuner = kt.RandomSearch(build_model, objective='val_mean_absolute_error',
max_trials=5, executions_per_trial=3,
directory='project', project_name='Air Quality Index')
tuner.search(X_train, y_train, epochs=5, validation_data=(X_test, y_test))
best_model = tuner.get_best_models(num_models=1)[0]RandomSearch vs GridSearchCV
| Keras Tuner RandomSearch | GridSearchCV | |
|---|---|---|
| Library | keras_tuner, for Keras models | scikit-learn, for scikit-learn models |
| How it picks | random combinations, max_trials of them | every combination of the grid |
| Scoring | a validation set, executions_per_trial runs averaged | k-fold cross-validation |
| Search space | can depend on itself (units per layer) | a fixed grid |
| Covered in | this lesson | Hyperparameter tuning with GridSearchCV |
Where you use Keras Tuner
- Choosing the layer sizes of a tabular ANN like the churn network instead of picking 11, 7 and 6 by hand.
- Tuning the learning rate and dropout rate together, with
hp.Choiceandhp.Float. - Searching CNN filters and kernel sizes with the same build_model pattern.
df.isna().sum() and drop or fill missing values before searching.Related
- Previous: Training and evaluating an ANN
- Next: Black box vs white box models
- See also: Hyperparameter tuning with GridSearchCV, MSE, MAE and Huber loss
- Reference: KerasTuner documentation
- Change
rng.integers(2, 21)torng.integers(2, 4)in the sampling run and compare the parameter counts. - Remove the
df = df.dropna()line in the data run and check which split the NaN target ends up in. - Count the parameters of the second-best trial (3 layers of 480, 480 and 512 units) with
n_params.
Little by little, you're building something great.