0% found this document useful (0 votes)
8 views16 pages

1996 Modularity

The document discusses a novel approach to building control systems for mobile robots using behavior-based robotics, which allows robots to select behaviors based on environmental stimuli. It presents experiments comparing different neural network architectures for controlling a robot tasked with cleaning an arena by picking up trash, highlighting that emergent modular architectures yield superior performance. The findings suggest that the interaction between modules and sensory-motor mappings is complex, indicating limitations in traditional engineering approaches to behavior-based robotics.

Uploaded by

Kam Bielawski
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views16 pages

1996 Modularity

The document discusses a novel approach to building control systems for mobile robots using behavior-based robotics, which allows robots to select behaviors based on environmental stimuli. It presents experiments comparing different neural network architectures for controlling a robot tasked with cleaning an arena by picking up trash, highlighting that emergent modular architectures yield superior performance. The findings suggest that the interaction between modules and sensory-motor mappings is complex, indicating limitations in traditional engineering approaches to behavior-based robotics.

Uploaded by

Kam Bielawski
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Institute of Psychology

C.N.R. - Rome

Using emergent modularity to develop control systems for


mobile robots

Stefano Nolfi
Institute of Psychology, National Research Council, Rome, Italy.
e-mail: stefano@[Link]

June 1996 (revised December 1996)

Technical Report 96-14

Department of Neural Systems and Artificial Life


15, Viale Marx
00137 - Rome - Italy
voice: 0039-6-86090231
fax: 0039-6-824737

To appear on: Adaptive Behavior, Special issue on "Complete agent learning in complex
environment"
Using emergent modularity to develop control systems for mobile robots
Stefano Nolfi
Institute of Psychology, National Research Council
15, Viale Marx - 00187 - Rome - Italy
voice: 0039-6-86090231
fax: 0039-6-824737
stefano@[Link]

Abstract

A new way of building control systems, known as behavior based robotics, has recently been proposed
to overcome the difficulties of the traditional AI approach to robotics. This new approach is based
upon the idea of providing the robot with a range of simple behaviors and letting the environment
determine which behavior should have control at any given time. We will present a set of experiments
in which neural networks with different architectures have been trained to control a mobile robot
designed to keep an arena clear by picking-up trash objects and releasing them outside the arena.
Controller weights are selected using a form of genetic algorithm and do not change during the
lifetime (i.e. no learning occurs). We will compare, in simulation and on a real robot, five different
network architectures and will show that a network which allows for fine-grained modularity achieves
significantly better performance. By comparing the functionality of each network module and its
interaction with a description of the simple behavior components, we will show that it is not possible
to find simple correlations; rather, module switching and interaction is correlated with low-level
sensory-motor mappings. This implies that the engineering-oriented approach to behavior-based
robotics might have serious limitations because it is difficult to know in advance the appropriate
mappings between behavior components and sensory-motor activity for complex tasks.

1. Introduction us “approach” or “avoid”, which are specified


from the observer’s point of view, and
A new way of building control systems, designing an action selection or coordination
known as behavior based robotics, has mechanism able to ensure that only the correct
recently been proposed to overcome the basic behavior has the control over the
difficulties of the traditional AI approach to actuators at the right time. Most of the time
robotics (Brooks, 1986). This new approach is both the modules of the controller
based upon the idea of providing the robot corresponding to the defined basic behaviors
with a range of simple behaviors and letting and the action selection mechanism are
the environment determine which behavior designed by the experimenter even if the
should have control at any given time. Despite design process is accomplished often
the central role of the environment however, incrementally and involves intensive testing
behavior based systems differ from purely and debugging. However, it has also been
reactive systems because “they can use shown that basic behaviors and action
different forms of internal representations and selection can be learned (Maes 1992,
perform computations on them in order to Mahadevan and Connell, 1992; Dorigo and
decide what effector action to take” (Mataric, Schnepf, 1993).
1992). We claim that, in order to really obtain
The design of a control system centered on simple and robust solutions, the process of
the behavior based approach usually involves breaking down the required behavior into sub-
breaking down the required behavior into a set components and of integrating them should be
of basic behaviors (also called reflexes), such accomplished taking into account the

1
“proximal” description of behavior (i.e. the of the observer in which high level terms such
sensory-motor loops responsible for the as “approach” or “attack” are used to describe
resulting behavior of an individual) and not the the result of a sequence of sensory-motor loops
“distal” description of behavior (i.e. a high (distal description of behavior) and a
level in which terms such as “approach” or description from the point of view of the agent’s
“avoid” are used to describe, from the sensory-motor system that accounts for how the
observer’s point of view, the result of a agent itself reacts in different sensory situations
sequence of sensory-motor loops). This can be (proximal description of behavior). The distal
accomplished, as we will show in this paper, description of a behavior is a function not only
by using learning or adaptation not only to of the controller determining how the agent
develop the module of the controller reacts to each possible sensory stimulus but
responsible for the basic behaviors and action also of the environment and of the agent’s
selection, but also to break the required sensory and motor apparatus. As a
behavior down into basic behaviors to be consequence, simple controllers may be able to
coordinated/selected. produce behaviors that, even if simple in a
To investigate this issue we will present a proximal description, may appear complex in a
set of experiments in which neural networks distal description.
with different architectures were trained to Another important characteristic of the
control a mobile robot designed to keep an behavior of mobile agents is that by actively
arena clear by picking-up trash objects and interacting with the external environment, for
releasing them outside the arena. The obtained example by moving, they are able to partially
results, validated on the real robot, show how determine the kind of stimuli they are exposed
the best performances are obtained with an to (see Parisi, Cecconi, Nolfi, 1990). A
emergent modular architecture (i.e. an response to a sensory state results in a new
architecture in which the different modules sensory state and in a new motor response that
responsible for the different sub-behaviors, the in its turn results in a new sensory state,
action selection mechanism, and the process of
forming a sensory-motor loop. This fact has
breaking down the required behavior into basic
several implications: (a) agents can produce
behaviors are the result of an adaptation
behavioral sequences without using any form of
process). Moreover, the analysis of the
memory; (b) agents can self-select favorable
successfully trained individuals shows that
sensory states and avoid unfavorable ones
there are no correspondence between neural
(Nolfi and Parisi, 1993). This also implies that
modules and distal description of behaviors
not all the sensory-motor states that constitute
and that in order to understand the role of
the proximal description of an agent’s behavior
modularity one should look at the proximal
description of behavior. have the same importance. Some sensory-motor
states may be extremely important (for example
1.1 Describing behaviors of agents that interact because they allow the agent to avoid falling
with an external environment into an undesirable behavioral loop), while
others may be completely irrelevant because
The main goal of research in robotics is they are never encountered (given the fact that
the attempt to provide a methodology for the agents themselves determine the successive
synthesizing control systems for physical sensory states) or because, whatever motor
robots able to produce complex behavior. response the agent produces, it in any case falls
However, the fact that simple mechanisms, in within a behavioral loop.
the appropriate environment, may exhibit
complex behaviors (Braitenberg, 1984) raises 2. The experimental framework
the problem of defining more clearly the terms
behavior and complexity (referred to Having at our disposal a Khepera robot
behavior). with the gripper module (see below) we
One way to overcome this problem is, as decided to try to develop a control system for a
proposed by Sharkey and Heemskerk (in robot with the task of keeping clear an arena
press), to distinguish two ways of describing surrounded by walls. The robot should look for
behavior: a description from the point of view "garbage", somehow grasp it, and take it out of

2
the arena. The task of cleaning the arena can on to the floor when pushed by the robot. Nest
be broken down into several sub-tasks: (a) position is marked by another cylinder
explore the environment, avoiding the walls; wrapped in pink paper. The robot uses a
(b) recognize a target object and to place the frontal color camera to identify the position of
body in a relative position so that it can be food cylinders and of the nest, using colors to
grasped; (c) pick up the target object; (d) discriminate. Moreover, the nest sensor uses
move toward the walls while avoiding other an odometer to get the approximate position of
target objects; (e) recognize a wall and place the nest when it is not visible.
the body in a relative position that allows the To build the controller the authors
object to be dropped out of the arena; (g) decomposed the target behavior into a
release the object. Moreover, these sub-tasks collection of simple behaviors (leave-nest, get-
can be broken down into smaller components. food, reach-nest, avoid-obstacles, coordinate-
For example (a) may be broken down into (a1) behaviors) and allocated a behavioral module
go forward when sensors are not activated; to each of them (behavioral modules have been
(a2) turn left at a given speed when right implemented using classifier systems). The
sensors are activated etc. However, as we want behavioral modules (with the exception of the
the complete solution to the task to emerge obstacle-avoidance module, which was pre-
through an evolutionary process we do not programmed) were trained separately and then
need to specify the requested behavior in detail frozen. The coordinator module was then
or to analyze the interference between basic trained to achieve the target behavior.
behaviors. However, they decided how to decompose the
Scheier and Pfeifer (1995) developed the target behavior into basic behaviors while we
control systems for a Khepera robot that want also this subdivision to be the result of a
performs a task very similar to the one training phase. As Colombetti, Dorigo, and
described in this paper (see also Scheier and Borghi note at the end of their paper: “In
Lambrinos, 1995). The environment is an order to relieve designers from part of their
arena surrounded by walls which contains burden, learning techniques might be extended
large and small pegs and a home base with a to other aspects of robot development, like the
light source attached to it. The robot has to architecture of the controller. This means that
bring the small pegs to the home base. It was the structure of behavioral modules should
decided to design by hand a set of modules emerge from the learning process, instead of
each corresponding to an elementary behavior being pre-designed.” (Colombetti, Dorigo, and
(move forward, turn toward objects, avoid Borghi, 1996).
obstacles, grasp, and bring to the nest). All that We decided to train controllers using an
is acquired during the training phase is the evolutionary method (Cliff, Harvey, and
tuning of the grasp behavior. The robot is pre- Husband, 1993; Nolfi, Floreano, Miglino, and
programmed to turn around pegs and the size Mondada, 1994; Mataric, and Cliff, in press)
of the pegs determines the way in which the and to conduct the training process in
angular velocity of the robot changes in time. simulation (for a description of the simulator
Reinforcement learning is used to associate the see Miglino, Lund, and Nolfi, 1995). The
vectors of angular velocities corresponding to control systems were then downloaded into the
small pegs with grasping behavior; in other robot and tested in the real environment. In
words to classify the two types of pegs. On the this section we will describe the robot and the
other hand, in this paper we want the entire environment, the architecture of the controller,
control system, including its organization into and the genetic algorithm used. In section 3 we
modules corresponding to basic behaviors, to will describe the results obtained and the
emerge during the training phase. characteristics and advantages of the emergent
Colombetti, Dorigo, and Borghi (1996) modular architecture described.
also studied a task similar to that described in
this paper. They trained a mobile robot based 2.1. The robot and the environment
on a commercial platform produced by
RoboSoft to collect food pieces and to store Khepera is a miniature mobile robot
them in a nest. Each piece of food is a developed at E.P.F.L. in Lausanne,
cylinder, wrapped in violet paper, which slides Switzerland (Mondada, Franzi, and Ienne,

3
1993). It has a circular shape with a diameter neural networks are resistant to noise, which is
of 55 mm, a height of 30 mm, and a weight of massively present in robot/environment
70g. It is supported by two wheels and two interactions and are potentially able of
small Teflon balls. The wheels are controlled generalizing their behavior to new situations;
by two DC motors with an incremental (b) it is important that the primitives
encoder (10 pulses per mm of advancement by manipulated by the evolutionary process
the robot), and they can move in both should be at the lowest possible level in order
directions. In addition, the robot is provided to avoid undesirable choices being made by
with a gripper module with two degrees of the human designer (Cliff, Harvey, and
freedom. The arm of the gripper can move Husband, 1993), and synaptic weights and
through any angle from vertical to horizontal neurons are sufficiently low level primitives;
while the gripper can assume only the open or (c) neural networks can easily exploit various
closed position. The robot is provided with form of learning during life-time, and this
eight infra-red proximity sensors (six sensors learning process may help and speed up the
are positioned on the front of the robot, and the evolutionary process (Ackley and Littman,
remaining two on the back), and an optical 1991; Nolfi, Elman and Parisi, 1994; Floreano
barrier sensor on the gripper capable of and Mondada, 1996).
detecting the presence of an object within the In order to assess the role of modularity,
gripper (the infra-red sensors on the back side and of emergent modularity in particular, we
of the robot and other available sensors were tried several different network architectures.
not used in the experiments described in this All architectures had 7 sensory neurons and 4
paper). motor neurons although they differed in their
A Motorola 68331 controller with 256 internal organization. The first 6 sensory
Kbytes of RAM and 512 Kbytes ROM handles neurons were used to encode the activation
all the input-output routines and can level of the corresponding 6 frontal sensors of
communicate via a serial port with a host Khepera and the seventh sensory neuron was
computer. Khepera was attached to the host used to encode the barrier light sensor on the
computer by means of a lightweight aerial gripper. On the motor side the four neurons
cable and specially designed rotating contacts. respectively coded for the speed of the left and
This configuration makes it possible to trace right motors and for the triggering of the
and record all important variables by "object pick-up" and "object release"
exploiting the storage capabilities of the host procedures.
computer, and at the same time provides The activation values of the infrared
electrical power without using time-consuming sensors (which can have 1024 different values
homing algorithms or large heavy-duty ranging from 0 to 1023) and of the activation
batteries. of the light-barrier sensor (which can have two
The environment was a rectangular arena values: 0 or 1023) were encoded in sensory
60x35 cm surrounded by walls containing 5 neurons as floating point values between 0.0
target objects. The walls were 3 cm in height, and 1.0. The logistic function was used to
made of wood, and covered with white paper. determine the activation of the motor neurons.
Target objects consisted of cylinders with a The activation of the first two motor neurons
diameter of 2.3 cm and a height of 3 cm. They controlling the left and right wheels was
were made of cardboard and covered with transformed into 21 different integer values
white paper. Targets were positioned randomly ranging from -10 to +10 (max. speed backward
inside the arena. and forward, respectively). The activation of
the third and fourth motor neurons controlling
2.2 The architecture of the controller the picking-up and releasing procedures,
respectively, were thresholded into two values
Like the majority of people who use (1 = trigger the corresponding procedure, 0 =
evolutionary methods to obtain control do not trigger the corresponding procedure).
systems for autonomous robots (Mataric and
Cliff, in press), we decided to implement the
controller using a neural network. This
decision was based on several reasons: (a)

4
C
A

D
B

E
Figure 1. The 5 different architectures used to evolve the controller: (a) a standard feedforward architecture;
(b) an architecture with an internal layer of hidden units; (c) a recurrent architecture; (d) a modular
architecture with two pre-designed modules; (e) an emergent modular architecture.

The simplest architecture used was a 2- unspecified number of previous sensory


layer feedforward neural network (see Figure stimuli (for a similar architecture see Elman,
1a). The second architecture was also a 1990). Finally we tried two modular neural
feedforward neural network but had an internal architectures (i.e. networks in which different
layer of four units (Figure 1b). We then tried a parts or modules had control in different
recurrent architecture in which the activation sensory-environmental situations). The first
level of two additional output units was copied modular architecture (Figure 1d) had two
back into two additional input units (Figure modules, of which the corresponding expected
1c). We chose this architecture because it behavior was pre-determined by the designer.
allows the network: (a) to determine which The first module (i.e. the sub-network on the
type of information to keep in memory, and left) had control when the robot gripper was
(b) to compress into a single pattern of empty, and was therefore dedicated to the
activation information coming from an ability to find a target, while avoiding walls,

5
recognize it, and pick it up correctly. The determined the motor output when the module
second module (i.e. the sub-network on the has control, the second output neuron
right) was in control when the gripper was (selector) competes with the selector neuron of
carrying a target and was therefore dedicated the other corresponding module to determine
to the ability to find a wall while avoiding which of the two modules has to take control.
other targets, stop in front of it and release the The activation of the sensors and the state
target. The partition of the required behavior of the motors were encoded every 100
into these two basic behaviors and into the milliseconds. However, when the activation
corresponding neural modules was of course level of the "object pick-up" or of the "object
arbitrary, although it seemed to be the most release" neurons reached a given threshold, a
reasonable one given that the robot was sequence of action occurred that possibly
expected to perform two very different required one or two seconds to complete (e.g.
behaviors depending on the state of the move a little further back, close the gripper,
gripper. move the arm up, for the object pick-up
The second modular architecture (see Figure procedure; move the arm down, open the
1e) was denoted as an “emergent modular gripper, and move the arm up again, for the
architecture” because it allows the required object release procedure).
behavior to be broken down into sub- It is important to note that the task chosen
components corresponding to different neural is particularly well suited to study the role of
modules, although it does not require the modularity because, as described above, the
designer to do such a partition in advance. The required behavior can be broken down into
number of available neural modules (in this several basic behaviors that may be
case two for each motor output), the implemented in different neural modules.
architecture of each module, and the Moreover, the task requires a controller able to
mechanisms that determine their interaction is produce very different motor responses for
pre-designed and fixed. However the number similar sensory states. Let us take the case of
of modules actually used by an individual, the the robot in front of a target, it should avoid or
combination of modules used each time step, approach it according to the presence or
and the weights of the modules themselves are absence of a target on the gripper (in the two
learned during the training phase and are cases the only difference is the state of 1
emergent. In particular, the sub-division of the sensor out of 7). Or else, let us take the case of
behavior into basic behavior corresponding to a robot in front of an object with an empty
different neural modules is emergent. This can gripper, it should avoid or approach the object
be accomplished because the neural structures according the type of the object; wall or target
responsible for the basic behaviors and for the (in the two cases the infrared sensors have only
selection mechanisms are represented slightly different activation values). Our
homogeneously. hypothesis is that a modular neural network,
This architecture had 16 output units, that can use different neural modules in
which, at every time step, gives 4 output different environmental situations, might have
values controlling the 4 previously described an advantage in learning to produce very
effectors. Four pairs of output neurons different motor responses for very similar
(represented by empty circles) coded for the sensory patterns with respect to a single,
speed of the left and right motors and for the uniformly connected, neural network.
triggering of the "object pick-up" and "object
release" procedures, respectively, and four 2.3. The Genetic Algorithm
pairs of selector neurons (represented by full
circles) determined which of the two To evolve neural controllers able to
competing output neurons had control over the perform the task described above we used a
corresponding robot's effector each time step form of genetic algorithm (Holland, 1975). For
(the competitor with the corresponding highly each network architecture, we began with 100
activated selector neuron gained control). Each randomly generated genotypes each
module was composed of two output neurons, representing a network with the corresponding
two corresponding biases, and 14 connections architecture and a different set of randomly
from sensory neurons. The first output neuron assigned connection weights. This is

6
Generation 0 (G0). G0 networks are allowed to In this paper we will concentrate on (b) and
"live" for 15 epochs, with each epoch (c). We plan to investigate (a) in the near
consisting of 200 actions (about 8 seconds in future.
the simulated environment using an IBM The genetic encoding scheme was a direct
RISC/6000 and about 300 seconds in the real one-to-one mapping. The encoding scheme is
environment). At the beginning of each epoch the way in which the phenotype (in this case
the robot and the target objects were randomly the connection weights of the neural network)
positioned in the arena. Epochs terminated is encoded in the genotype (the representation
after 200 actions or after the first object had according to which the genetic algorithm
been correctly released. At the end of their life, operates). One-to-one mapping is the simplest
individual robots were allowed to reproduce. encoding scheme where one and only one
However, only the 20 individuals which had 'gene' corresponds to each phenotypical
accumulated the most fitness in the course of character. In our case, to each connection
their life reproduced (agamically) by weight and bias corresponded to a sequence of
generating 5 copies of their neural networks. 8 bits for the genotype which had a total length
These 20x5=100 new robots constituted the of: (A) 256, (B) 416, (C) 480, (D) 480, and (E)
next generation (G1). Mutations were 1024 bits in the 5 different architectures
introduced in the copying process, resulting in described. (For more complex encoding
possible changes of the connection weights. schemes also allowing evolution of the neural
Mutations were obtained by substituting 2% of architecture, see Cliff, Harvey and Husband,
randomly selected bits with a new randomly 1993; Nolfi, Miglino, and Parisi, 1994; Grau,
selected value (as a consequence, about 1% of 1995).
the bits were actually changed). We tried lower Individual networks were scored by
and higher mutation rates in our experiment. counting the number of objects correctly
This was selected because it gave the best released outside the arena. However, in order
results overall. The process was repeated for to facilitate the emergence of the ability to
1000 generations. achieve the task, individuals were also scored
We also ran a set of simulations in which (even if with a much lower reward) for their
we used both mutation and crossover. In this ability to pick up targets. In addition, it was
case a random single point crossover was found important to expose the robots to useful
performed with a given probability (0.1, 0.5, training experiences (i.e. to artificially increase
1.0). Reproducing individuals were obtained the number of times when the robot, while
by crossing over one of the 20 individuals with carrying a target object, encountered another
the highest fitness score and one randomly target) in order to force the evolutionary
selected individual. However, we did not process to select individuals able to avoid
obtain better performance with respect to targets when the gripper was full. This was
simulations without crossover (for this reason accomplished by artificially positioning a new
we shall present the result obtained using only target object in the frontal area of the robot
mutations). This may be due to the fact that each time it picked up a target during
the crossover points were randomly chosen. evolutionary training (see also Nolfi, in press).
Restricting the crossover points so as to Without this manipulation of the learning
preserve the organization of the network may experiences, evolved individuals were unable
produce better results (Montana, and Davis, to avoid targets while carrying an object. We
1989). This may be particularly true for first tried to introduce a penalty term in the
architectures D and E which can be divided fitness function for individuals unable to avoid
into neural modules. targets when their gripper was full. However
Gene duplication and elimination could alteration of the learning experiences proved
also be considered. These operators can be much more effective. The real problem, in fact,
particularly effective in the case of architecture is that such cases rarely occurred during
E in order to allow evolution to select the best training and as a consequence, without altering
number of competing neural modules for each the learning experiences, there was too little
motor output. Modularity can be realized at evolutionary pressure to select individuals able
different levels: (a) the genetic level; (b) the to perform the right behavior.
nervous system level; (c) the behavioral level.

7
3. Results other four conditions at generation 999 (p <
0.05).
We ran 10 simulations for each of the 5
different architectures described above. Each 12
E
B
simulation started with populations of 100 CD
9
networks with randomly assigned connection

successful epochs
weights and lasted 1000 generations (about 10 6
A

hours using a standard IBM RISC/6000). In


the following section we will compare the 3

results obtained for individuals with different


0
architectures, and later we will analyze how 0 100 200 300 400 500 600 700 800 900 1000
generations
modularity is used in emergent modular
architectures.
Figure 2. Number of epochs (out of 15) in which
individuals with different architectures correctly
3.1. Results obtained with different
picked up and then released a target object outside
architectures
the arena through out generations. Each curve
represents the average of the best individuals in 10
If we measure the average number of different simulations. Data smoothed by
epochs (out of 15) in which individuals calculating rolling averages over preceding and
correctly pick up and then release a target succeeding 3 generations.
outside the arena for simulations with different
architectures we can see how in all conditions By downloading the best controllers of
an ability to accomplish the correct sequence generation 999 (for 10 replications of the
of behaviors evolves (see Figure 2). Note that simulation) into the robot and testing them in
epochs terminated after 200 actions or after the the real environment for 5000 cycles for their
first object had been correctly released. As a ability to clean up the arena by removing 5
consequence the max. number of targets that randomly placed target objects, we can see that
can be released outside the arena is equal to architecture (E) clearly outperforms all other
the number of epochs. However, evolved architectures (see Figure 3). The best
individuals with different architectures vary in individuals of 7 (out of 10) with the emergent
the performance achieved at the end of the modular architecture were capable of cleaning
evolutionary training and in the time needed to the arena without displaying any incorrect
reach plateau level performance. If we look at behavior while only 1 or 2 individuals (out of
performance of generation 999 we can see how 10) with other architectures were capable of
all types of architectures have reached high accomplishing the task.
performances with the exception of the simple
feed-forward architecture (A). If we look at 7
performance throughout generations we can 6
see how the emergent modular architectures,
successful individuals

5
after few generations, start to outperform all 4
the other architectures maintaining a difference
3
until generation 500 (this result is even more
2
meaningful if one consider that architecture
1
(E), by requiring a longer genotype with
0
respect to the other architectures, also implies A B C D E

that the genetic algorithm has to search a larger


space). A oneway analysis of variance of Figure 3. Number of evolved individuals for each
performance in the five different conditions control architectures capable of correctly picking
was performed each 100 generations. The up and then releasing outside the arena the 5
results show that performance in condition (E) targets objects within 5000 cycles without
is significantly higher than in the other four displaying any incorrect behavior (e.g. crashing
conditions at generation 199 and performance into walls, trying to grasp a wall, or trying to
in condition (A) is significantly lower than the release a target over another target). A,B,C,D,E

8
indicate the five different architectures described competing for the control of the right motor
in Figure 1. are both used in all the phases that can be
described as distal sub-behaviors: when the
These results show that the emergent gripper is empty and the robot has to look for a
modular architecture (E) enables the target (i.e. when sensor LB is off); when the
evolutionary process to find a correct solution gripper is carrying a target and the robot has to
to the task earlier than other architectures and look for a wall (i.e. when sensor LB is on);
in particular earlier than the hand-crafted when the robot perceives something and has to
modular architecture (D). Moreover results disambiguate between walls and targets (i.e.
show how the emergent modular architecture when the W/T graph shows the upper or
allows the evolutionary process to select more bottom line); when the robot does not perceive
robust solutions to the task, i.e. controllers anything (i.e. when the ‘W/T’ graph does not
which showed only a limited loss of show any line); when the robot is approaching
performance when transferred into the real a target (i.e. when sensor LB is off and the
robot. perceived object is a target); when the robot is
approaching a wall (i.e. when sensor LB is on
3.2. How emergent modular neural networks and the perceived object is a wall); when the
work robot is avoiding a target (i.e. when sensor LB
is on and the perceived object is a target);
The first thing we want to know about our when the robot is avoiding a wall (i.e. when
evolved individuals with the emergent the sensor LB is off and the perceived object is
modular architecture is: can we find a a wall).
correspondence between distal description of Similar results can be obtained by
behaviors and modules? In other words do we analyzing the other evolved individuals. When
find that the evolved individuals use different many alternative neural modules are involved
modules in different environmental situations it becomes difficult to understand what is
(e.g. when they have to pick up a target, when going on. However, the general picture
they have to release a target, when they have remains the same: neural modules or a
to disambiguate a sensory pattern, when they combination of neural modules does not
have to avoid a target etc.)? appear to be responsible for single distal sub-
The answer to this question is no. behaviors. On the contrary each sub-behavior
Figure 4 represents the behavior of a is the result of the contribution of different
typical evolved individual. As we said in the neural modules.
previous section, individuals with emergent In order to understand the function of
modular architecture have two different modules in these evolved individuals we
modules for each of the four motor outputs and should abandon the analysis of distal behaviors
therefore can use up to 16 different and concentrate on proximal behaviors (i.e. the
combination of neural modules. However, the way in which individuals respond to different
evolved individual described in Figure 4, i.e. sensory stimuli). In the case of our
one of the most successful, uses only a single robot/environment framework, the number of
module to control the left motor, the pick-up stimuli (intended as the combination of all
procedures, and the release procedure (LM, possible sensor states) to which individuals
PU, and RL) and it uses both neural modules may be exposed is infinite. However, they can
only for the right motor (RM). For an analysis be reduced to a finite number by classifying
of other individuals see below. What is similar stimuli together.
interesting to note is that those two modules

9
Figure 4. The top part of the figure represents the behavior of a typical evolved individual in its environment.
Lines represent walls, empty and full circles represent the original and the final position of the target objects
respectively, the trace on the terrain represents the trajectory of the robot. The bottom part of the figure
represents the type of object currently perceived, the state of the motor, and the state of the sensors throughout
time for 500 cycles respectively. The ‘W/T’ graph shows whether the robot is currently perceiving a wall (top
line), a target (bottom line), or nothing (no line). The ‘LM’, ‘RM,’ ‘PU’, and ‘RL’ graphs show the state of the
motors (left and right motors, pick-up and release procedures, respectively). For each motor, in the top part of
the graph the activation state is indicated (after the arbitration between component modules has been
performed by the selector neurons) and in the bottom part which of the two competing neural modules has
control is indicated (the thickness of the line at the bottom indicates whether the first or the second module has
control: thin line segment = module 1; thick line segment = module 2). The graphs ‘I0’ to ‘I5’ show the state
of the 6 infrared sensors. Finally, the ‘LB’ graph shows the state of the light-barrier sensor. The activation
state of sensor and motor neurons is represented by the height with respect to the baseline (in the case of motor
neurons the activation state of the output neurons of the module that currently have the control is shown).

If we assume that stimuli perceived by the threshold of 2 degrees and 2 millimeters and
robot do not change significantly when the considering that infra-red sensors are unable to
robot modifies its position with respect to a detect objects at a distance of over 40
perceived object by turning left or right less millimeters, we obtain 180x20=3600 different
than 2 degrees or by moving forth or back less stimuli for different relative positions of the
the 2 millimeters (i.e. the same grain used to robot with respect to an object (i.e. 180
build the simulator through world samples (see different orientations from 0 to 359 degrees
Miglino, Lund, and Nolfi, 1995)) we can multiplied by 20 different distances from 0 to
reduce the infinite number of different input 40 millimeters). In addition, because stimuli
stimuli to a manageable number. By using a also vary according to the type of object

10
perceived (wall or target) and the state of the each of the 3600x4 different environmental
light-barrier sensor (on or off), we have a total stimulation, in addition to the state of the four
of 3600x4=14400 possible different sensory effectors, the corresponding combination of
stimuli. Figure 5 shows the proximal neural modules that obtain the control is
description of the behavior of the same indicated.
evolved individual represented in Figure 4. For

Figure 5. Proximal representation of the behavior of an evolved individual. The rectangular maps labeled ‘left
motor’, ‘right motor’, ‘pick up’, and ‘release’ represent the activation state of the four actuators (after the
arbitration between component modules has been performed by the selector neurons) for 180 different
o
orientations over 360 and 20 different distances (from 0 to 40 mm) with respect to a perceived object. For
graphic reasons the activation states of the motor neurons are divided into two classes, positive and negative
speeds (represented with black and white color respectively), although the left and right motors can take 20
different speeds. The ‘right motor winner’ maps represent which of the two modules controlling the right
motor is in charge for each combination of orientations and distance with respect to a perceived object (the
two modules are represented in white and black, respectively). The group of 5 maps is replicated for 4
different conditions: perceived object is a wall and no object is in the gripper (top left maps); perceived object
is a target and no object is in the gripper (bottom left maps); perceived object is a wall and an object is in the
gripper (top right maps); perceived object is a target and an object is in the gripper (bottom right maps).

If we analyze in which environmental the “area” in which the module represented in


situations the two different combinations of black obtains control varies significantly
neural modules obtain control we can see how, according to whether the perceived object is a
in the case of the individual represented in wall or a target (it is much larger in the case of
Figure 5, the neural module shown in black a wall) and, although much less significantly,
obtains the control when the perceived object whether the robot has an object in the gripper
o
is within a given angle (from about -100 to or not (it is larger in the second case).
o
100 ) and a given distance (about 30mm) with As can be seen, there is almost a one to
respect to the robot. However, the extension of one correspondence between the combination

11
of neural modules that have control (right architecture will increase with increasing task
motor winner) and the speed of the right complexity.
motor, which is the only motor affected by the
alternation of the two competing neural 3.3 Why there is no correspondence between
modules in this individual. The speed of the evolved modules and distal description of
right motor is negative when the black neural behaviors.
module is activated and positive otherwise.
Therefore the weights that determine which of There are at least two reasons that can
the two neural modules has control have the explain why there is no correspondence
main responsibility in determining the speed of between distal description of behaviors and
the right motor. However, the weights of the neural modules in evolved individuals. First,
two neural modules are responsible for one should consider that different distal
differentiating the speed of the right motor in a behaviors usually require behavioral responses
very important environmental situation (when that differ only partially. Think of the
both angle and distance are close to 0) following situations: object in the gripper or
depending on the presence or not of an object not. If the robot has an object in the gripper it
in the gripper (see the left and right ‘right should avoid targets and approach walls and
motor’ maps in the bottom of Figure 5). viceversa when it has the gripper empty it
By observing the close descriptions of should approach targets and avoid walls.
behavior of other evolved individuals However the perceived object can be correctly
(obtained by replicating the simulation) it classified only from a small number of relative
appears that different combinations of modules positions (see Nolfi, 1996). Therefore, in all
are used to produce different motor responses the other cases, there is no reason to have
for similar sensory stimuli when necessary. different behaviors involving different neural
This happens most of the time, as in the case modules. Conversely having two different
of the individual described in Figure 5, when neural modules for the cases “object in the
the robot has an object on its frontal side and gripper or not”, as in architecture D described
must decide whether to approach or avoid it or above, requires the additional cost of learning
whether to try to pick it up or not. Individuals the same behavior in all cases in which there is
with the other architecture described appear no reason to have different motor responses.
less able to produce sharp discontinuities in Alternation of different neural modules is
behavior. required only when the agent is expected to
The need to produce very different motor produce different motor responses for similar
responses for very similar sensory stimuli is input patterns, and this happens only in some
related to the complexity of the task. This type cases in different distal behaviors.
of architecture may consequently be expected One second reason is that the division of a
to scale up more easily than other non-modular requested behavior into a set of basic behavior
architectures. We have not jet applied this that correspond to distal description often
architecture to more complex tasks than the requires a very complex behavioral selection
one described in this paper. However we have system. Let us take the basic behaviors:
observed that while homogeneous avoiding a target, avoiding a wall, approaching
architectures are perfectly able to solve simpler a target, and approaching a wall. All these
tasks like the ability to explore an arena behaviors are probably easy to implement (or
surrounded by walls (Nolfi, Floreano, Miglino, to learn); however they require a selection
and Mondada, 1994) or the ability to mechanism that, in order to be able to select
recognize, approach and to remain close to the right module at the right time, should
target objects while avoiding walls in an always be able to correctly classify walls and
environment identical to that described in this target and this is a complex task that can be
paper (Nolfi, 1996), the emergent modular accomplished only in some environmental
neural network outperformed other circumstances. In our emergent modular
architectures in garbage collecting task. architecture, because the neural modules and
Therefore, one can hypothesize that the the selection mechanism are represented
advantage of the emergent modular homogeneously and evolve at the same time,

12
solutions in which both the components are By evolving controllers for mobile robots
kept as simple as possible are selected. for the purpose of performing a non trivial task
An interesting property of the emergent and by comparing the results obtained using
modular network is that it does not require this architecture with other neural architectures
specification of the number of modules we showed that the emergent modular
corresponding to basic behaviors into which architecture outperforms all other architectures
the desired behavior should be broken down. It and in particular the modular architecture in
is only necessary to specify the max. number which the break-down of the required behavior
of available neural modules for each motor into basic behaviors is hand-crafted.
which then determines the max. number of The analysis of the evolved individuals
different combinations of neural modules that with the emergent modular architecture showed
can be exploited. In the present paper we used that modularity and action selection can be
an architecture with 8 neural modules (two useful in tasks, like that presented in this paper,
neural modules for each of the four motor in which very different motor responses should
functions) which can produce up to 16 be produced for similar sensory patterns. In
different combinations of neural modules. other words modularity appears useful in tasks
However, only about 6 neural modules and which require complex behavior from the point
about 5 combinations of neural modules were of view of proximal description. Moreover, the
used, on average, by evolved individuals. The analysis of our results showed that there is no
number of neural modules actually selected by correspondence between evolved modules and
evolution can be expected to be, on average, distal description of behaviors. In fact, the
proportional to the complexity of the task. sequence of sensory-motor loops that can be
This is an important property of the described as basic behaviors from the
architecture. In fact if we assume that the best observer’s point of view are the result of the
way to break down a behavior into basic contribution of different neural modules in
behaviors corresponding to different neural evolved individuals with the emergent modular
modules is to take into account the close architecture.
description of behaviors, and we also assume The fact that the process of breaking down
that there is a complex mapping between close the required behavior into basic behaviors
and distal description of behavior we will should take into account the proximal
conclude that: as the designer will be unable to description of behavior itself can impose
specify the best way to break down the serious limitations on the engineering-oriented
required behavior into basic behaviors, he will approach to behavior-based robotics. This is
also be unable to specify how to select the
because there is a complex mapping between
number of basic behaviors. Nor will the
distal and proximal description of behavior and
designer able to decide how to combine
therefore we cannot expect the experimenter,
different neural modules for each time step.
who has direct access only to the distal
description of behavior, to have a correct
4. Conclusions
picture of the corresponding proximal
We have presented an architecture, known descriptions, excepted for trivial cases. Even
as emergent modular architecture, in which for the use of learning in the development of the
each output function two or more alternative modules responsible for the basic behavior
neural modules compete for control and two or and/or of the mechanisms responsible for action
more other corresponding neural modules selection might suffice. Hand-crafting the
determine which competitor gains control. process of breaking down the required behavior
This architecture, by using a uniform into basic components leads to constraints that
representation for modules responsible for may limit the adaptation process to the borders
basic behavior and mechanisms responsible for set by the experimenter.
behavior selection, allows not only the control Our claim that behavioral modules should
structures responsible for basic behaviors and be allocated by considering the proximal
behavior selection but also the break down of description and not the distal description of
the target behavior into basic behaviors to be behavior, as is usually done, may only be true
obtained through an adaptation process. (or particularly true) for neurocontrollers

13
where a homogeneous set of weighted sums Gruau, F. (1995). Automatic definition of modular
are often requested to account for sharp neural networks. Adaptive Behavior, 2, 151-
discontinuities in behavior. However, similar 183.
impressions have been reported by other Holland, J. H. (1975). Adaptation in Natural and
researchers following different approaches. Artificial Systems, Ann Arbor, Mich.,
University of Michigan Press.
Mahadevan and Connel, for example, who
Maes, P. (1992). Learning behavior networks from
developed a box-pushing controller for an experience, in: F. J. Varela, P. Bourgine (eds.),
autonomous robot using a subsumption Toward a Practice of Autonomous Systems:
architecture wrote “Obviously, there may be Proceedings of the First European Conference
several ways of decomposing a given task, and on Artificial Life, Cambridge, Mass, MIT
coming up with a good decomposition is a Press/Bradford Books.
nontrivial problem” (Mahadevan, and Connel; Mahadevan, S., & Connell, J. (1992). Automatic
1992, p.363). programming of behavior-based robots using
reinforcement learning. Artificial Intelligence,
Acknowledgment 55, 311-365.
Mataric, M. J. (1992). Behavior-based control:
This research has been granted by the Main properties and implications. Proceedings
Coordinated Project on real Time Computing in of the IEEE International Conference on
Robotics and Autonomation, Workshop on
Real World of C.N.R, Italy. The author thanks
Architectures for Intelligent Control Systems,
the anonymous referees for valuable Nice, France.
suggestions on the manuscript. Mataric, M. J., & Cliff, D. (in press). Challenges
in evolving controllers for physical robots, in
References “Evolutionary Robotics”, special issue of
Robotics and Autonomous Systems.
Ackley, D. H., & Littman, M. L. (1991). Miglino, O., Lund, H. H., & Nolfi, S. (1995).
Interactions between learning and evolution, in: Evolving mobile robots in simulated and real
C. G. Langton, J. D. Farmer, S. Rasmussen, C. environments. Artificial Life, (2) 4, 417-434.
E. Taylor (eds.), Artificial Life II, Reading, Mondada, F., Franzi, E., & Ienne, P. (1993).
Mass., Addison-Wesley. Mobile Robot miniaturisation: A tool for
Braitenberg, V. (1984). Vehicles: experiments in investigation in control algorithms, in:
synthetic psychology, Cambridge, MA: MIT Proceedings of the Third International
Press. Symposium on Experimental Robotics, Kyoto,
Brooks, R. A. (1986). A roboust layered control Japan.
system for a mobile robot, IEEE Journal of Montana, D. J., & Davis, L. (1989). Training
Robotics and Autonomation, 2, 14-23. feedforward neural networks using genetic
Cliff, D. T., Harvey, I., & Husbands, P. (1993). algorithms, in: Proceedings of Eleventh Joint
Explorations in Evolutionary Robotics. Conference on Artificial Intelligence, Vol 1,
Adaptive Behavior, 2, 73-110. Palo Alto, CA: Kaufmann.
Colombetti, M., Dorigo, M., & Borghi G. (1996). Nolfi, S. (1996). Adaptation as a more powerful
Behavior analysis and training. A methodology tool than decomposition and integration, in T.
for behavior engineering. IEEE Transactions Fogarty and G. Venturini (eds), Proceedings of
on Systems, Man, and Cybernetics - Part B, the workshop on Evolutionary computing and
(26) 1, 29-41 Machine Learning, 13th International
Dorigo, M., & Schnepf, U. (1993). Genetic-based Conference on Machine Learning, Bari.
machine learning and behaviour based robotics: Nolfi, S. (in press). Evolving non-trivial behaviors
a new synthesis. IEEE Transaction on Systems, on real robots: a garbage collecting robot.
Man, and Cybernetics, (23) 1, 141-154. Robotics and Autonomous Systems.
Elman, J. L. (1990) Finding structure in time, Nolfi, S., Elman, J.L., & Parisi, D. (1994).
Cognitive Science, 14, 179-211 Learning and Evolution in Neural Networks.
Floreano, D., & Mondada, F. (1996). Evolution of Adaptive Behavior, 1, 5-28.
plastic neurocontrollers for situated agents, in: Nolfi, S., Floreano, D., Miglino, O., & Mondada,
P. Maes, M. Mataric, J-A. Meyer, J. Pollack, F. (1994). How to evolve autonomous robots:
& S. Wilson. (eds.), From Animals to Animats different approaches in evolutionary robotics,
IV, Cambridge, MA: MIT Press. in: R.A. Brooks and P. Maes (eds.),
Proceedings of fourth International

14
Conference on Artificial Life, Cambridge,
Mass, MIT Press.
Nolfi, S., Miglino, O., & Parisi, D. (1994).
Phenotypic Plasticity in Evolving Neural
Networks, in: D. P. Gaussier and J-D. Nicoud
(eds.) Proceedings of the Intl. Conf. From
Perception to Action, Los Alamitos, CA: IEEE
Press.
Nolfi, S., & Parisi, D. (1993). Self-selection of
input stimuli for improving performance. In: G.
A. Bekey (ed.) Neural Networks and Robotics,
Kluwer Academic Publisher.
Parisi, D., Cecconi, F., & Nolfi, S. (1990). Econets:
Neural networks that learn in an environment.
Network,1,149-168.
Sharkey, N. E., & Heemskerk, N. H. (in press).
The neural mind and the robot, in A. J. Browne
(Ed.) Current Perspective in Neural
Computing, IOP press.
Scheier C., & Pfeifer, R. (1995). Classification as
sensory-motor coordination: A case study on
autonomous agents, in: F. Moran, A. Moreno,
J.J. Merelo, P. Chacon (Eds.) Advances in
Artificial Life: Proceedings of the Third
European Conference on Artificial Life,
Springer Verlag.
Scheier, C., & Lambrinos, D. (1995). Adaptive
classification in autonomous agents. Technical
Report, AILab, Computer Science Department,
University of Zurich.

15

You might also like