Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2014 Jul 24;9(7):e103143.
doi: 10.1371/journal.pone.0103143. eCollection 2014.

Reward maximization justifies the transition from sensory selection at childhood to sensory integration at adulthood

Affiliations

Reward maximization justifies the transition from sensory selection at childhood to sensory integration at adulthood

Pedram Daee et al. PLoS One. .

Erratum in

  • PLoS One. 2014;9(12):e115926

Abstract

In a multisensory task, human adults integrate information from different sensory modalities--behaviorally in an optimal Bayesian fashion--while children mostly rely on a single sensor modality for decision making. The reason behind this change of behavior over age and the process behind learning the required statistics for optimal integration are still unclear and have not been justified by the conventional Bayesian modeling. We propose an interactive multisensory learning framework without making any prior assumptions about the sensory models. In this framework, learning in every modality and in their joint space is done in parallel using a single-step reinforcement learning method. A simple statistical test on confidence intervals on the mean of reward distributions is used to select the most informative source of information among the individual modalities and the joint space. Analyses of the method and the simulation results on a multimodal localization task show that the learning system autonomously starts with sensory selection and gradually switches to sensory integration. This is because, relying more on modalities--i.e. selection--at early learning steps (childhood) is more rewarding than favoring decisions learned in the joint space since, smaller state-space in modalities results in faster learning in every individual modality. In contrast, after gaining sufficient experiences (adulthood), the quality of learning in the joint space matures while learning in modalities suffers from insufficient accuracy due to perceptual aliasing. It results in tighter confidence interval for the joint space and consequently causes a smooth shift from selection to integration. It suggests that sensory selection and integration are emergent behavior and both are outputs of a single reward maximization process; i.e. the transition is not a preprogrammed phenomenon.

PubMed Disclaimer

Conflict of interest statement

Competing Interests: The authors have declared that no competing interests exist.

Figures

Figure 1
Figure 1. Different types of perceptual aliasing in subspaces.
[Image: see text] represents the observation set of the ith sensor for i = 1, 2. [Image: see text] is the state set and A = {○,□,Δ} is the action set of the agent. [Image: see text] and [Image: see text] are the best and the worst actions in the given state, respectively. Accumulated experience in [Image: see text] is a perfect generalization for [Image: see text] and [Image: see text], since these two states have the same optimal policy and [Image: see text] is common between them. In contrast, accumulated experience in [Image: see text] is garbage information because functionally different states are mapped to the same observation. The situation for [Image: see text] and [Image: see text] is a little different. Only for the best action in [Image: see text] and the worst action in [Image: see text] we have the generalization, however, for the other action this is not the case.
Figure 2
Figure 2. A schematic overview of the proposed framework for multisensory learning and decision making.
s = (o1,o2,…,ok) is the perceptual input, [Image: see text] is the current reading of the ith sensor, and [Image: see text] is the learning block of the ith sensor. For each action and based on the previously received rewards, each learning block calculates a confidence interval ([Image: see text]) on the mean of the reward distribution corresponding to the given observation and action pair. The proposed Generalization Test (G Test), tests the generalization ability of the individual source against the joint space. In case that an individual source passes the G Test, its confidence interval will be considered in the decision making phase. In decision making phase, an appropriate action based on the given intervals will be selected which considers the exploration and exploitation trade-off.
Figure 3
Figure 3. Stimulus and observations by the auditory ([Image: see text]) and the visual ([Image: see text]) sensors.
Observations are based on Gaussian noise models. Variances control the reliability of each sensor.
Figure 4
Figure 4. Performance and behavior of the method in the localization task.
All graphs are results of averaging over 20 independent runs and passing a moving average window with size 500. (A) Average reward for all agents. For the proposed methods (MOS and LUS), we used Table 3, employing bound (1) with [Image: see text] for calculating confidence intervals. The rival methods employ the UCB1 policy on the individual sensors and on the joint space. (B) Average acceptance rate (1–rejection rate) of the individual sensors in the proposed method (MOS). (C) The average dominancy percentage of each source in decision making (MOS). In the first half of learning steps, vision is the dominant sensor while the agent prefers the integrated sensory data in the rest of learning steps.
Figure 5
Figure 5. Performance of the method (MOS) in response to an unexpected change in the environment.
At time step [Image: see text] the visual sensor fails and its variance changes to the highest possible value. All graphs are results of averaging over 10 independent runs and passing a moving average window with size 500. (A) Average acceptance rate ([Image: see text]) of the individual sensors. (B) The average dominancy percentage of each source in decision making (MOS). After failure of the visual sensor, the method detects this change and relies on the auditory sensor for decision making.
Figure 6
Figure 6. Impact of [Image: see text].
We used four different values (0.05, 0.25, 0.45, 0.80) for [Image: see text] from being conservative to liberal in terms of confidence. All graphs are results of averaging over 10 independent runs and passing a moving average window with size 500. (A) Average acceptance rate (1–rejection rate) of the individual sensors in the proposed method (MOS). The upper/lower ribbon for each value of [Image: see text] represents visual/auditory sensor. By increasing [Image: see text], the test becomes harder for the individual sensors to pass. (B) The average dominancy percentage of each source in decision making (MOS). For each value of [Image: see text], the ascending ribbon represents integration and the two descending ribbons represent selection of visual and auditory sensors. Increasing [Image: see text] results in earlier cross of the ascending and the descending ribbons; i.e. earlier switch from selection to integration.
Figure 7
Figure 7. Performance and behavior of the method in response to an unreliable sensor.
All graphs are results of averaging over 20 independent runs and passing a moving average window with size 1000. (A) Average reward for all agents. For the proposed method (MOS), we used Table 3, employing bound (1) with [Image: see text] for calculating confidence intervals. The rival methods employ the UCB1 policy on the individual sensors and on the joint space. (B) Average acceptance rate (1–rejection rate) of the individual sensors in the proposed method (MOS). (C) The average dominancy percentage of each source in decision making (MOS). Due to unreliability of the noise sensor, it takes longer for learning in the integrated states to mature and, therefore, dominancy of the visual sensor is prolonged.
Figure 8
Figure 8. Dominancy of subspaces over time.
The average dominancy percentage of different combination of sensors in decision making (LUS). Subspaces including the unreliable source have been filtered. Furthermore, dependency on the integration of reliable sensors increases over time.

References

    1. Ernst MO, Banks MS (2002) Humans integrate visual and haptic information in a statistically optimal fashion. Nature 415: 429–433. - PubMed
    1. Alais D, Burr D (2004) The Ventriloquist Effect Results from Near-Optimal Bimodal Integration. Current Biology 14: 257–262. - PubMed
    1. Burr D, Gori M (2012) Multisensory Integration Develops Late in Humans. In: Murray MM, Wallace MT, editors. The Neural Bases of Multisensory Processes. Boca Raton (FL): CRC Press. - PubMed
    1. Gori M, Del Viva M, Sandini G, Burr DC (2008) Young Children Do Not Integrate Visual and Haptic Form Information. Current Biology 18: 694–698. - PubMed
    1. Nardini M, Jones P, Bedford R, Braddick O (2008) Development of Cue Integration in Human Navigation. Current Biology 18: 689–693. - PubMed