Panasonic Patent | Audio signal processing method, recording medium, and audio signal processing device
Patent: Audio signal processing method, recording medium, and audio signal processing device
Publication Number: 20260230772
Publication Date: 2026-08-06
Assignee: Panasonic Intellectual Property Corporation Of America
Abstract
An audio signal processing method executed by an audio signal processing device includes: obtaining an audio signal having an attribute indicating an indirect sound; calculating a first sound volume based on a sound volume of the indirect sound when arriving at a listening position, using: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal; calculating a second sound volume based on a sound volume of a direct sound, associated with the indirect sound, when arriving at the listening position; selecting whether to output an output signal based on the audio signal, using: a sound volume ratio between the second and first sound volumes; and a time difference between the direct and indirect sounds; and outputting the output signal, when outputting the output signal is selected.
Claims
1.An audio signal processing method executed by an audio signal processing device, the audio signal processing method comprising:obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; calculating a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on:a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
2.The audio signal processing method according to claim 1, whereinin the calculating of the second sound volume, the second sound volume is calculated based on a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the direct sound.
3.The audio signal processing method according to claim 1, whereinthe frequency characteristic indicating the auditory sensitivity is a frequency characteristic indicating sound volume sensitivity of the listener.
4.The audio signal processing method according to claim 3, whereinthe frequency characteristic indicating the sound volume sensitivity is a frequency characteristic based on an inverse of an equal loudness contour.
5.The audio signal processing method according to claim 3, whereinthe frequency characteristic indicating the sound volume sensitivity is an A-weighting characteristic.
6.The audio signal processing method according to claim 3, whereinthe frequency characteristic indicating the sound volume sensitivity is an inverse characteristic of a frequency characteristic of a minimum audible angle of a sound source position.
7.An audio signal processing method executed by an audio signal processing device, the audio signal processing method comprising:obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating a reflected sound that is a sound resulting from a direct sound being reflected by a reflector; selecting whether to output an output signal that is based on the audio signal obtained, based on:a sound volume ratio between: a sound volume of the reflected sound at a time at which the reflected sound arrives at a listening position that is a position at which a listener is present; and a sound volume of the direct sound at a time at which the direct sound arrives at the listening position, the reflected sound being indicated by the audio signal obtained; a time difference between when the direct sound arrives and when the reflected sound arrives; and a reflection coefficient characteristic amount determined based on a reflection coefficient of the reflector; and outputting the output signal, when outputting the output signal is selected.
8.The audio signal processing method according to claim 7, whereinthe reflection coefficient characteristic amount is a characteristic amount indicating a degree of flatness of a frequency characteristic of a reflection coefficient.
9.An audio signal processing method executed by an audio signal processing device, the audio signal processing method comprising:obtaining an audio signal; calculating a predetermined sound volume, the predetermined sound volume being based on a sound volume of a sound at a time at which the sound arrives at a listening position that is a position at which a listener is present, the sound being a sound indicated by the audio signal obtained, the calculating being based on:a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound; and the audio signal obtained; selecting whether to output an output signal that is based on the audio signal obtained, based on the predetermined sound volume calculated; and outputting the output signal, when outputting the output signal is selected.
10.An audio signal processing method executed by an audio signal processing device, the audio signal processing method comprising:obtaining an audio signal including attribute information identifying an attribute of the audio signal; calculating a first sound volume, the first sound volume being based on a sound volume of an indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on:a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the audio signal for which the attribute includes information indicating the indirect sound, the attribute being identified by the attribute information included in the audio signal obtained; and gain information of the indirect sound; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound, the calculating of the second sound volume being based on:a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the direct sound; and gain information of the direct sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
11.The audio signal processing method according to claim 10, whereinthe gain characteristic of each predetermined frequency bandwidth related to the indirect sound is stored as the attribute information related to the indirect sound, the gain information related to the indirect sound is stored as the attribute information of the indirect sound, the gain characteristic of each predetermined frequency bandwidth related to the direct sound is stored as the attribute information related to the direct sound, and the gain information related to the direct sound is stored as the attribute information of the direct sound.
12.A non-transitory computer-readable recording medium having recorded thereon a computer program for causing a computer to execute the audio signal processing method according to claim 1.
13.An audio signal processing device comprising:an obtainer that obtains an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; a first calculator that calculates a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the first calculator calculating the first sound volume based on:a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; a second calculator that calculates a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; a selection processor that selects whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and a reproducer that outputs the output signal, when outputting the output signal is selected.
Description
CROSS REFERENCE TO RELATED APPLICATIONS
This is a continuation application of PCT International Application No. PCT/JP2024/035621 filed on Oct. 4, 2024, designating the United States of America, which is based on and claims priority of U.S. Provisional Patent Application No. 63/542,854 filed on Oct. 6, 2023. The entire disclosures of the above-identified applications, including the specifications, drawings and claims are incorporated herein by reference in their entirety.
FIELD
The present disclosure relates to an audio signal processing method and the like.
BACKGROUND
In recent years, the spread of products and services that utilize extended reality (ER) (may be also expressed as “XR”) including virtual reality (VR), augmented reality (AR), and mixed reality (MR) has advanced. Accompanying this, there has been growing demand for audio signal processing technologies that provide listeners with immersive audio that, in a virtual space or a real-world space, assigns acoustic effects that are generated in accordance with the environment of the space to sounds emitted from a virtual sound source.
Note that “listener” can also be expressed as “user”. Furthermore, Patent Literature (PTL) 1, PTL 2, PTL 3, and Non Patent Literature (NPL) 1 disclose techniques that relate to the audio signal processing method and the like of the present disclosure.
CITATION LIST
Patent Literature
PTL 1: Japanese Patent No. 6288100PTL 2: Japanese Unexamined Patent Application Publication No. 2019-22049PTL 3: WO Publication No. 2021/180938
Non Patent Literature
NPL 1: B. C. J. Moore, “An Introduction to the Psychology of Hearing”, Seishin Shobo, 1994 Apr. 20, Chapter 6: Space Perception, p. 225.
SUMMARY
Technical Problem
Incidentally, in the technique disclosed in PTL 1, it may be difficult to appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Accordingly, the present disclosure provides an audio signal processing method and the like that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Solution to Problem
An audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; calculating a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Furthermore, an audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating a reflected sound that is a sound resulting from a direct sound being reflected by a reflector; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between: a sound volume of the reflected sound at a time at which the reflected sound arrives at a listening position that is a position at which a listener is present; and a sound volume of the direct sound at a time at which the direct sound arrives at the listening position, the reflected sound being indicated by the audio signal obtained; a time difference between when the direct sound arrives and when the reflected sound arrives; and a reflection coefficient characteristic amount determined based on a reflection coefficient of the reflector; and outputting the output signal, when outputting the output signal is selected.
Furthermore, an audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal; calculating a predetermined sound volume, the predetermined sound volume being based on a sound volume of a sound at a time at which the sound arrives at a listening position that is a position at which a listener is present, the sound being a sound indicated by the audio signal obtained, the calculating being based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound; and the audio signal obtained; selecting whether to output an output signal that is based on the audio signal obtained, based on the predetermined sound volume calculated; and outputting the output signal, when outputting the output signal is selected.
Furthermore, an audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal; calculating a first sound volume, the first sound volume being based on a sound volume of an indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the audio signal for which the attribute includes information indicating the indirect sound, the attribute being identified by the attribute information included in the audio signal obtained; and gain information of the indirect sound; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound, the calculating of the second sound volume being based on: a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the direct sound; and gain information of the direct sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Furthermore, a recording medium according to one aspect of the present disclosure is a non-transitory computer-readable recording medium having recorded thereon a computer program for causing a computer to execute the audio signal processing method described above.
Furthermore, an audio signal processing device according to one aspect of the present disclosure is an audio signal processing device including: an obtainer that obtains an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; a first calculator that calculates a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the first calculator calculating the first sound volume based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; a second calculator that calculates a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; a selection processor that selects whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and a reproducer that outputs the output signal, when outputting the output signal is selected.
Note that these comprehensive or specific aspects may be implemented as a system, a device, a method, an integrated circuit, a computer program, or a non-transitory computer-readable recording medium such as a CD-ROM, or may be implemented as any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
Advantageous Effects
The audio signal processing method and the like according to the one aspect of the present disclosure make it possible to appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
BRIEF DESCRIPTION OF DRAWINGS
These and other advantages and features will become apparent from the following description thereof taken in conjunction with the accompanying Drawings, by way of non-limiting examples of embodiments disclosed herein.
FIG. 1 is a diagram illustrating one example of a direct sound and reflected sounds generated in a sound space.
FIG. 2 is a diagram illustrating an example of a three-dimensional sound reproduction system according to Embodiment 1.
FIG. 3A is a block diagram illustrating a configuration example of an encoding device according to Embodiment 1.
FIG. 3B is a block diagram illustrating a configuration example of a decoding device according to Embodiment 1.
FIG. 3C is a block diagram illustrating another configuration example of an encoding device according to Embodiment 1.
FIG. 3D is a block diagram illustrating another configuration example of a decoding device according to Embodiment 1.
FIG. 4A is a block diagram illustrating a configuration example of a decoder according to Embodiment 1.
FIG. 4B is a block diagram illustrating another configuration example of a decoder according to Embodiment 1.
FIG. 5 is a diagram illustrating an example of a physical configuration of an audio signal processing device according to Embodiment 1.
FIG. 6 is a diagram illustrating an example of a physical configuration of an encoding device according to Embodiment 1.
FIG. 7 is a block diagram illustrating a configuration example of a renderer according to Embodiment 1.
FIG. 8 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 1.
FIG. 9 is a diagram illustrating a comparatively distant positional relationship between a listener and an obstacle object.
FIG. 10 is a diagram illustrating a comparatively close positional relationship between a listener and an obstacle object.
FIG. 11 is a diagram illustrating relationships between time differences between direct sounds and reflected sounds, and threshold values.
FIG. 12A is a diagram illustrating a part of an example of a method for setting threshold value data.
FIG. 12B is a diagram illustrating a part of an example of a method for setting threshold value data.
FIG. 12C is a diagram illustrating a part of an example of a method for setting threshold value data.
FIG. 13 is a diagram illustrating an example of a threshold value setting method.
FIG. 14 is a flowchart illustrating an example of selection processing.
FIG. 15 is a diagram illustrating relationships between directions of direct sounds, directions of reflected sounds, time differences, and threshold values.
FIG. 16 is a diagram illustrating relationships between angular differences, time differences, and threshold values.
FIG. 17 is a block diagram illustrating another configuration example of a renderer.
FIG. 18 is a flowchart illustrating another example of selection processing.
FIG. 19 is a flowchart illustrating yet another example of selection processing.
FIG. 20 is a flowchart illustrating a first variation of operations of an audio signal processing device according to Embodiment 1.
FIG. 21 is a flowchart illustrating a second variation of operations of an audio signal processing device according to Embodiment 1.
FIG. 22 is a diagram illustrating an arrangement example of an avatar, a sound source object, and an obstacle object.
FIG. 23 is a flowchart illustrating yet another example of selection processing.
FIG. 24 is a block diagram illustrating a configuration example for a renderer to perform pipeline processing.
FIG. 25 is a diagram illustrating transmission and diffraction of a sound.
FIG. 26 is a diagram illustrating an example of the positional relationship between a listener and an obstacle object, according to Embodiment 1.
FIG. 27 is a diagram illustrating another example of the positional relationship between a listener and an obstacle object, according to Embodiment 1.
FIG. 28 is an example of an echo detection limit threshold value according to Embodiment 1.
FIG. 29 is a block diagram illustrating a configuration example of a renderer according to Embodiment 2.
FIG. 30 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 2.
FIG. 31 is a flowchart illustrating an operation example of the selection processing according to Embodiment 2.
FIG. 32 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth related to a reflected sound according to Embodiment 2.
FIG. 33A is a diagram illustrating a table showing a frequency characteristic indicating auditory sensitivity according to Embodiment 2.
FIG. 33B is a diagram illustrating the frequency characteristic indicating the auditory sensitivity according to Embodiment 2.
FIG. 34 is a diagram illustrating the gain characteristic related to the reflected sound according to Embodiment 2, the frequency characteristic indicating the auditory sensitivity (A-weighting characteristic), and a first correction characteristic.
FIG. 35 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth related to a direct sound according to Embodiment 2.
FIG. 36 is a diagram illustrating the gain characteristic related to the direct sound, the frequency characteristic indicating auditory sensitivity (A-weighting characteristic), and a second correction characteristic, according to Embodiment 2.
FIG. 37 is a diagram illustrating the inverse characteristic of an equal loudness contour, which is another, first example of the frequency characteristic indicating the sound volume sensitivity of the listener according to Embodiment 2.
FIG. 38 is a diagram illustrating the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position, which is another, second example of the frequency characteristic indicating the sound volume sensitivity of the listener according to Embodiment 2.
FIG. 39 is a block diagram illustrating a configuration example of a renderer according to Embodiment 3.
FIG. 40 is a diagram illustrating the impact of a reflection coefficient characteristic amount on the first threshold value, according to Embodiment 3.
FIG. 41 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 3.
FIG. 42 is a block diagram illustrating a configuration example of a renderer according to Embodiment 4.
FIG. 43 is a graph illustrating threshold value data indicating a second threshold value according to Embodiment 4.
FIG. 44 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 4.
FIG. 45 is a flowchart illustrating an operation example of the selection processing according to Embodiment 4.
DESCRIPTION OF EMBODIMENTS
(Underlying Knowledge Forming Basis of the Present Disclosure)
To date, investigations have been made into audio signal processing technologies that provide listeners with immersive audio by, in a virtual space or a real-world space, assigning acoustic effects that are generated in accordance with the environment of the space to sounds emitted from a virtual sound source.
PTL 1 discloses such an audio signal processing technique. More specifically, PTL 1 discloses a technique for detecting the importance of audio signals (voice signals) and not outputting audio signals for which the detected importance is low. By thus not outputting audio signals for which the importance is low, there are expectations for the audio signal processing technique to appropriately reduce the amount of computation and the computational load.
Incidentally, in a sound space (a virtual space or a real-world space), reflected sound is sometimes important.
FIG. 1 is a diagram illustrating one example of a direct sound and reflected sounds generated in a sound space. In acoustic processing in which characteristics of a virtual space are expressed by a sound, it is effective to reproduce not only direct sounds, but also reflected sounds in order to express the size of the space, the material of the walls, and the like, as well as to allow for accurately grasping the location of the sound source (the positioning of the sound image).
For example, when a sound is heard in a rectangular parallelepiped room such as that in FIG. 1, six primary reflected sounds, corresponding to the six walls, are generated for one sound source. Reproducing these reflected sounds provides a clue for appropriate understanding of the space and the sound image. Furthermore, for each reflected sound, a secondary reflected sound is generated by a surface other than the reflection surface that generated that reflected sound. These reflected sounds are also effective sensory clues.
However, even when consideration is given no further than to secondary reflection, one direct sound and 36 (6+6×5) reflected sounds are generated for one sound source. Thus, 37 sound rays are generated, and processing these sound rays requires a significant amount of computation.
Furthermore, in applied products in recent years for which metaverses are imagined, such as virtual meetings, virtual shopping, virtual concerts, and the like, a plurality of sound sources are present out of necessity, whereby an even greater amount of computation is required.
Moreover, the listener hearing the sounds in a virtual space uses headphones or VR goggles. In order to provide three-dimensional sound to such a listener, binaural processing that assigns a sound pressure ratio and a phase difference between the two ears and reproduces the direction of arrival and distance sensation of the sounds is performed on each sound ray. Thus, if an attempt were made to reproduce every reflected sound that is generated, the amount of computation would become immense.
On the other hand, in light of convenience, a small storage battery is sometimes used as the battery for the VR goggles worn by the listener who experiences the virtual space. Lessening the computational load resulting from the above-described processing makes it possible to further extend the life of the storage battery. To this end, the number of sound rays, which are emitted on a scale of hundreds, is desirably reduced, within a scope at which grasping the space and the positioning of the sounds is not harmed.
Furthermore, in a system that reproduces acoustics, a degree of freedom such as 6DoF (6 degrees of freedom) or the like may be allowed with respect to the position (in other words, the listening position, which is the position where the listener is present) and orientation of the listener. In this case, the positional relationship between the listener, the sound sources, and the objects that reflect sounds cannot be fixed until the time of reproduction (the time of rendering). For this reason, the reflected sounds as well cannot be fixed until the time of reproduction. Thus, it is difficult to determine the reflected sounds to be processed beforehand.
Therefore, during reproduction, appropriately selecting and outputting (reproducing) one or more reflected sounds, from a plurality of reflected sounds that are generated in a sound space, that are to be processed or are not to be processed is useful in appropriately reducing the amount of computation and the computational load.
It should be noted that controlling whether to select a sound corresponds to determining whether to select the sound, and more specifically corresponds to determining whether to select and output (reproduce) the sound. Furthermore, selecting a sound may be selecting the sound as a sound to be processed, or may be selecting the sound as a sound that is not to be processed.
Incidentally, in PTL 1, the importance of the audio signal, and more specifically the importance of the direct sound that the audio signal represents, is detected, but the importance of reflected sound is not considered. Therefore, when indirect sounds such as reflected sounds are generated, as illustrated in FIG. 1, the amount of computation and the computational load increase; in other words, it may be difficult to appropriately reduce the amount of computation and the computational load.
Furthermore, in conventional techniques including the technique disclosed in PTL 1, whether to output an audio signal is selected without taking the auditory sensitivity of the listener into consideration. When such an audio signal is output and the listener hears the sound indicated by that audio signal, the listener hears a sound that differs from his/her own auditory perception, causing a sense of incongruence.
Therefore, there is a need for an audio signal processing method and the like that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration, in a sound space.
Accordingly, an audio signal processing method according to a first aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; calculating a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Whether to output the output signal based on the audio signal indicating the indirect sound is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Furthermore, the first sound volume is calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the first sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
An audio signal processing method according to a second aspect of the present disclosure is the audio signal processing method according to the first aspect, wherein in the calculating of the second sound volume, the second sound volume is calculated based on a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the direct sound.
The second sound volume is thus calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the second sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into further consideration.
An audio signal processing method according to a third aspect of the present disclosure is the audio signal processing method according to the first or second aspect, wherein the frequency characteristic indicating the auditory sensitivity is a frequency characteristic indicating sound volume sensitivity of the listener.
This makes it possible to realize an audio signal processing method that enables using the frequency characteristic indicating the sound volume sensitivity of the listener as the frequency characteristic indicating the auditory sensitivity.
An audio signal processing method according to a fourth aspect of the present disclosure is the audio signal processing method according to the third aspect, wherein the frequency characteristic indicating the sound volume sensitivity is a frequency characteristic based on an inverse of an equal loudness contour.
This makes it possible to realize an audio signal processing method that enables using the frequency characteristic based on the inverse of the equal loudness contour as the frequency characteristic indicating the sound volume sensitivity.
An audio signal processing method according to a fifth aspect of the present disclosure is the audio signal processing method according to the third aspect, wherein the frequency characteristic indicating the sound volume sensitivity is an A-weighting characteristic.
This makes it possible to realize an audio signal processing method that enables using the A-weighting characteristic as the frequency characteristic indicating the sound volume sensitivity.
An audio signal processing method according to a sixth aspect of the present disclosure is the audio signal processing method according to the third aspect, wherein the frequency characteristic indicating the sound volume sensitivity is an inverse characteristic of a frequency characteristic of a minimum audible angle of a sound source position.
This makes it possible to realize an audio signal processing method that enables using the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position as the frequency characteristic indicating the sound volume sensitivity.
An audio signal processing method according to a seventh aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating a reflected sound that is a sound resulting from a direct sound being reflected by a reflector; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between: a sound volume of the reflected sound at a time at which the reflected sound arrives at a listening position that is a position at which a listener is present; and a sound volume of the direct sound at a time at which the direct sound arrives at the listening position, the reflected sound being indicated by the audio signal obtained; a time difference between when the direct sound arrives and when the reflected sound arrives; and a reflection coefficient characteristic amount determined based on a reflection coefficient of the reflector; and outputting the output signal, when outputting the output signal is selected.
Whether to output the output signal that is based on the audio signal indicating the reflected sound is thus selected based on the above-described sound volume ratio and the above-described time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Here, attention is directed to the precedence effect. The precedence effect is said to occur when the frequency spectrum of a leading sound (for example, a direct sound) approximates the frequency spectrum of a lagging sound (for example, a reflected sound). The frequency spectrum of reflected sound varies according to the reflection coefficient of the reflector. Therefore, whether the frequency spectrum of a direct sound approximates the frequency spectrum of a reflected sound varies in accordance with the reflection coefficient characteristic amount.
For example, when the reflection coefficient characteristic amount has a certain value, the frequency spectrum of the direct sound and the frequency spectrum of the reflected sound approximate each other, making it more likely for the precedence effect to occur. In this case, selection may be performed such that the output signal is less likely to be output. Furthermore, for example, when the reflection coefficient characteristic amount has another certain value, the frequency spectrum of the direct sound and the frequency spectrum of the reflected sound do not approximate each other, making it less likely for the precedence effect to occur. In this case, selection may be performed such that the output signal is more likely to be output.
That is, selecting whether to output the output signal based on the reflection coefficient characteristic amount is equivalent to selecting whether to output the output signal while taking the precedence effect into consideration. Since the precedence effect is an example of an auditory sensitivity characteristic, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
An audio signal processing method according to an eighth aspect of the present disclosure is the audio signal processing method according to the seventh aspect, wherein the reflection coefficient characteristic amount is a characteristic amount indicating a degree of flatness of a frequency characteristic of a reflection coefficient.
This makes it possible to realize an audio signal processing method that enables using, as the reflection coefficient characteristic amount, the characteristic amount indicating the degree of flatness of the frequency characteristic of the reflection coefficient.
An audio signal processing method according to a ninth aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal; calculating a predetermined sound volume, the predetermined sound volume being based on a sound volume of a sound at a time at which the sound arrives at a listening position that is a position at which a listener is present, the sound being a sound indicated by the audio signal obtained, the calculating being based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound; and the audio signal obtained; selecting whether to output an output signal that is based on the audio signal obtained, based on the predetermined sound volume calculated; and outputting the output signal, when outputting the output signal is selected.
The predetermined sound volume is thus calculated while taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal that is based on the audio signal representing the sound is selected, based on the predetermined sound volume calculated. That is, whether to output the output signal that is based on the audio signal is appropriately selected while taking the auditory sensitivity into consideration. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
An audio signal processing method according to a tenth aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal; calculating a first sound volume, the first sound volume being based on a sound volume of an indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the audio signal for which the attribute includes information indicating the indirect sound, the attribute being identified by the attribute information included in the audio signal obtained; and gain information of the indirect sound; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound, the calculating of the second sound volume being based on: a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the direct sound; and gain information of the direct sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Whether to output the output signal based on the audio signal indicating the indirect sound is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
An audio signal processing method according to an eleventh aspect of the present disclosure is the audio signal processing method according to the tenth aspect, wherein the gain characteristic of each predetermined frequency bandwidth related to the indirect sound is stored as the attribute information related to the indirect sound, the gain information related to the indirect sound is stored as the attribute information of the indirect sound, the gain characteristic of each predetermined frequency bandwidth related to the direct sound is stored as the attribute information related to the direct sound, and the gain information related to the direct sound is stored as the attribute information of the direct sound.
This makes it possible to realize an audio signal processing method in which various information is stored as the attribute information.
A recording medium according to a twelfth aspect of the present disclosure is a non-transitory computer-readable recording medium having recorded thereon a computer program for causing a computer to execute the audio signal processing method according to any one of the first to eleventh aspects.
This makes it possible for a computer to execute the above-described audio signal processing method, according to the computer program.
An audio signal processing device according to a thirteenth aspect of the present disclosure is an audio signal processing device including: an obtainer that obtains an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; a first calculator that calculates a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the first calculator calculating the first sound volume based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; a second calculator that calculates a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; a selection processor that selects whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and a reproducer that outputs the output signal, when outputting the output signal is selected.
Whether to output the output signal based on the audio signal indicating the indirect sound is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Furthermore, the first sound volume is calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the first sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing device that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Embodiment 1
(Example of Three-Dimensional Sound Reproduction System)
FIG. 2 is a diagram illustrating an example of three-dimensional sound reproduction system 1000. Specifically, FIG. 2 illustrates three-dimensional sound reproduction system 1000, which is an example of a system to which acoustic processing or decoding processing of the present disclosure can be applied. Three-dimensional sound is also expressed as immersive audio. Three-dimensional sound reproduction system 1000 includes audio signal processing device 1001 and audio presentation device 1002.
Audio signal processing device 1001, which is also expressed as an acoustic processing device, applies acoustic processing to an audio signal emitted from a virtual sound source and generates an acoustic-processed audio signal to be presented to the listener. The audio signal is not limited to voices, and is acceptable as long as it is an audible sound. Acoustic processing is, for example, signal processing applied to an audio signal in order to reproduce one or more effects that a sound receives between when the sound is emitted from a sound source and when the sound arrives at the listener.
Audio signal processing device 1001 performs acoustic processing based on spatial information that describes the main factors for bringing about the above-described effects. Spatial information encompasses, for example: information that indicates the location of a sound source, a listener, and objects in the vicinity; information that indicates the shape of a space; parameters regarding sound propagation; and the like. Audio signal processing device 1001 is, for example, a PC (personal computer), a smartphone, a tablet, a game console, or the like.
An acoustic-processed signal is presented from audio presentation device 1002 to the listener. Audio presentation device 1002 is connected to audio signal processing device 1001 via wireless or wired communication. The acoustic-processed audio signal generated by audio signal processing device 1001 is transmitted to audio presentation device 1002 via wireless or wired communication.
When audio presentation device 1002 includes a plurality of devices such as, for example, a device for the right ear and a device for the left ear, or the like, the plurality of devices present sound in synchronization by means of communication between the plurality of devices or communication between each of the plurality of devices and audio signal processing device 1001. Audio presentation device 1002 is, for example, headphones, earphones, or a head-mounted display worn on the head of the listener, surround speakers including a plurality of fixed speakers, or the like.
Note that three-dimensional sound reproduction system 1000 may be used in combination with an image presentation device or a stereoscopic image presentation device that visually provides an ER experience that includes AR/VR. For example, a space handled by spatial information is a virtual space in which the positions of sound sources, the listener, and objects in the space are virtual positions of virtual sound sources, a virtual listener, and virtual objects in a virtual space. The space can also be expressed as a sound space. Furthermore, the spatial information can also be expressed as sound space information.
Furthermore, FIG. 2 illustrates a system configuration example in which audio signal processing device 1001 and audio presentation device 1002 are separate devices, but three-dimensional sound reproduction system 1000 to which the acoustic processing method (the audio signal processing method) or the decoding method of the present disclosure can be applied is not limited to the configuration in FIG. 2. For example, audio signal processing device 1001 may be included in audio presentation device 1002, and audio presentation device 1002 may perform both acoustic processing and sound presentation.
Furthermore, audio signal processing device 1001 and audio presentation device 1002 may, in a shared manner, perform the acoustic processing described in the present disclosure. Furthermore, a server connected to audio signal processing device 1001 or audio presentation device 1002 over a network may perform a part or all of the acoustic processing described in the present disclosure.
Furthermore, audio signal processing device 1001 may perform the acoustic processing by decoding a bitstream that has been generated by encoding at least a part of data of the audio signal and the spatial information used in the acoustic processing. Thus, audio signal processing device 1001 may be expressed as a decoding device.
(Example of Encoding Device)
FIG. 3A is a block diagram illustrating a configuration example of encoding device 1100. Specifically, FIG. 3A illustrates the configuration of encoding device 1100, which is an example of the encoding device of the present disclosure.
Input data 1101 is data to be encoded, and includes spatial information and/or an audio signal to be inputted into encoder 1102. Details regarding the spatial information will be described later.
Encoder 1102 encodes input data 1101 to generate encoded data 1103. Encoded data 1103 is, for example, a bitstream generated by means of encoding processing.
Memory 1104 stores encoded data 1103. Memory 1104 may be, for example, a hard disk or an SSD (solid-state drive), or may be another type of memory.
Note that in the above description, a bitstream generated by means of encoding processing was given as an example of encoded data 1103 stored in memory 1104, but encoded data 1103 may be data other than a bitstream. For example, encoding device 1100 may store, in memory 1104, converted data generated by converting the bitstream into a predetermined data format. The converted data may be, for example, a file or multiplexed stream that corresponds to one or more bitstreams.
Here, the file is a file having a file format of, for example, ISO base media file format (ISOBMFF) or the like. Furthermore, encoded data 1103 may be in the form of a plurality of packets generated by splitting the above-described bitstream or file.
For example, the bitstream generated by encoder 1102 may be converted to data that is different from the bitstream. In this case, encoding device 1100 may include a converter, not illustrated, and the converter may perform conversion processing, or conversion processing may be performed by a central processing unit (CPU) that is an example of a processor, described later.
(Example of Decoding Device)
FIG. 3B is a block diagram illustrating a configuration example of decoding device 1110. Specifically, FIG. 3B illustrates the configuration of decoding device 1110, which is an example of the decoding device of the present disclosure.
Memory 1114 stores, for example, the same data as encoded data 1103 generated by encoding device 1100. The stored data is read from memory 1114 and inputted into decoder 1112 as input data 1113. Input data 1113 is, for example, a bitstream that is to be decoded. Memory 1114 may be, for example, a hard disk or an SSD, or may be another type of memory.
Note that decoding device 1110 may not directly input, to decoder 1112, data read from memory 1114 as input data 1113, and may instead convert the data read and then input the converted data to decoder 1112 as input data 1113. The data before conversion may be, for example, multiplexed data that includes one or more bitstreams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF or the like.
Furthermore, the data before conversion may be a plurality of packets generated by splitting the above-described bitstream or file. Data that is different from the bitstream may be read from memory 1114 and then converted into a bitstream. In this case, decoding device 1110 may include a converter, not illustrated, and the converter may perform conversion processing, or conversion processing may be performed by a CPU that is an example of a processor, described later.
Decoder 1112 decodes input data 1113 to generate audio signal 1111 that indicates audio to be presented to the listener.
(Other Example of Encoding Device)
FIG. 3C is a block diagram illustrating a configuration example of another encoding device. Specifically, FIG. 3C illustrates the configuration of encoding device 1120, which is another example of the encoding device of the present disclosure. In FIG. 3C, constituent elements that are the same as the constituent elements in FIG. 3A have been given the same reference signs as in FIG. 3A, and description of these constituent elements is omitted.
Encoding device 1100 stores encoded data 1103 in memory 1104. On the other hand, encoding device 1120 is different from encoding device 1100 in the respect that encoding device 1120 includes transmitter 1121 that transmits encoded data 1103 externally.
Transmitter 1121 transmits, to a different device or server, transmission signal 1122 that is generated based on data converted from encoded data 1103 or encoded data 1103 to a different file format. The data used in generating transmission signal 1122 is, for example, the bitstream, multiplexed data, file, or packet described in relation to encoding device 1100.
(Other Example of Decoding Device)
FIG. 3D is a block diagram illustrating another configuration example of a decoding device. Specifically, FIG. 3D illustrates the configuration of decoding device 1130, which is another example of the decoding device of the present disclosure. In FIG. 3D, constituent elements that are the same as the constituent elements in FIG. 3B have been given the same reference signs as in FIG. 3B, and description of these constituent elements is omitted.
Decoding device 1110 reads input data 1113 from memory 1114. On the other hand, decoding device 1130 is different from decoding device 1110 in the respect that decoding device 1130 includes receiver 1131, which receives input data 1113 from an external source.
Receiver 1131 receives reception signal 1132 to obtain reception data, and outputs input data 1113 to be inputted into decoder 1112. The reception data may be the same as input data 1113 inputted into decoder 1112, or may be data in a data format that is different from that of input data 1113.
When the data format of the reception data is different from the data format of input data 1113, receiver 1131 may convert the reception data into input data 1113. Alternatively, a converter or a CPU, each not illustrated, of decoding device 1130 may convert the reception data into input data 1113. The reception data is, for example, the bitstream, multiplexed data, file, or packet described in relation to encoding device 1120.
(Example of Decoder)
FIG. 4A is a block diagram illustrating a configuration example of decoder 1200. Specifically, FIG. 4A illustrates the configuration of decoder 1200, which is an example of decoder 1112 in FIG. 3B and FIG. 3D.
Input data 1113 is an encoded bitstream, and includes encoded audio data that is an audio signal that has been encoded, and metadata used in acoustic processing.
Spatial information manager 1201 obtains the metadata included in input data 1113 and analyzes the metadata. The metadata includes information that describes the main factors that act on the sounds arranged in the sound space. Spatial information manager 1201 manages the spatial information that is obtained by analyzing the metadata and is used in the acoustic processing, and provides the spatial information to renderer 1203.
Note that in the present disclosure, the information used in the acoustic processing is expressed as spatial information, but another expression may be used. For example, the information used in the acoustic processing may be expressed as sound space information, or may be expressed as scene information. Furthermore, when the information used in the acoustic processing changes over time, the spatial information inputted into renderer 1203 may be information expressed as a spatial state, a sound space state, a scene state, or the like.
Furthermore, the spatial information may be managed for each sound space or for each scene. For example, when each of a plurality of mutually differing rooms is expressed as a virtual space, the plurality of rooms may be managed as a plurality of scenes that mutually differ. Furthermore, spatial information may be managed such that even the same room is managed as a different scene in accordance with the expressed state.
Thus, a plurality of items of spatial information may be managed with respect to a plurality of sound spaces or a plurality of scenes. In management of a plurality of items of spatial information, an identifier that identifies each item of the plurality of items of spatial information may be assigned to the spatial information.
The spatial information data may be included in a bitstream that is an example of input data 1113. Alternatively, the bitstream may include an identifier of the spatial information, and the spatial information data may be obtained from an information source other than the bitstream. Specifically, when the bitstream includes only the identifier of the spatial information, in the rendering, the spatial information data stored in the memory inside the device or in an external server may be obtained as input data 1113, using the identifier of the spatial information.
Note that the information managed by spatial information manager 1201 is not limited to information included in the bitstream. For example, input data 1113 may include, as data not included in the bitstream, data that indicates the characteristics and structure of a space obtained from a VR or AR software application or server.
Furthermore, input data 1113 may include data that indicates the characteristics, position, and/or the like of the listener or an object. Moreover, input data 1113 may include information on the position of the listener, obtained using a sensor included in a terminal including a decoding device (1110, 1130), or may include information that indicates the position of the terminal, estimated based on information obtained using the sensor.
In other words, spatial information manager 1201 may communicate with an external system or server to obtain spatial information and the position of the listener (in other words, the listening position). Furthermore, spatial information manager 1201 may obtain clock synchronization information from an external system and perform processing to perform synchronization with a clock in renderer 1203.
Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real-world space or a virtual space that corresponds to a real-world space, i.e., an AR space or an MR space. Furthermore, the virtual space may be expressed as a sound field or a sound space. Moreover, the information indicating position in the above description may be information on coordinates or the like that indicate a position in a space, may be information that indicates a relative position with respect to a predetermined reference position, or may be information that indicates movement or acceleration of a position in a space.
Audio data decoder 1202 decodes encoded audio data included in input data 1113 to obtain an audio signal.
The encoded audio data obtained by three-dimensional sound reproduction system 1000 is, for example, a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO/IEC 23008-3). Note that MPEG-H 3D Audio is merely an example of an encoding method that can be used when generating the encoded audio data included in the bitstream. The encoded audio data may be a bitstream encoded by another encoding method.
For example, the encoding method may be a lossy codec such as MPEG-1 Audio Layer III (MP3), Advanced Audio Coding (AAC), Windows Media Audio (WMA), Audio Codec 3 (AC3), Vorbis, or the like. Alternatively, the encoding method may be a lossless codec such as Apple Lossless Audio Codec (ALAC), Free Lossless Audio Codec (FLAC), or the like.
Alternatively, any encoding method other than the above-described may be used. For example, PCM data may be a type of the encoded audio data. In this case, when, for example, the quantization bit rate of the PCM data is N, the decoding processing may be processing in which the N-bit binary number is converted into a numerical format (for example, floating-point format) that can be processed by renderer 1203.
Renderer 1203 obtains the audio signal and the spatial information, applies acoustic processing to the audio signal using the spatial information, and outputs an acoustic-processed audio signal (audio signal 1111).
Before starting the rendering, spatial information manager 1201 reads the metadata of the input signal, detects rendering items such as objects and sounds specified by the spatial information, and transmits the rendering items to renderer 1203. After the start of rendering, spatial information manager 1201 grasps the change over time of the spatial information and the position of the listener, and updates and manages the spatial information. Then, the updated spatial information is transmitted to renderer 1203.
Renderer 1203 generates and outputs an acoustic processing-added audio signal based on the audio signal included in input data 1113 and the spatial information received from spatial information manager 1201.
Update processing of the spatial information and output processing of the acoustic processing-added audio signal may be performed in the same thread. Furthermore, processing may be allocated to spatial information manager 1201 and renderer 1203 in mutually independent threads. When spatial information manager 1201 and renderer 1203 perform the update processing of the spatial information and the output processing of the acoustic processing-added audio signal in different threads, the thread activation frequency may be set separately, or processing may be performed in parallel.
When spatial information manager 1201 and renderer 1203 perform processing in independent threads that are different, renderer 1203 can be preferentially allotted computation resources. This makes it possible to safely execute sound output processing for which even a slight delay is unallowable, i.e., for which a popping noise would be generated with a delay of even one sample (0.02 msec).
At this time, allotment of computation resources to spatial information manager 1201 is limited. However, since compared to the output processing of the audio signal, updating of the spatial information is processing that is infrequent (for example, processing such as updating the orientation of a listener's face), it is not necessary for the output processing of the audio signal to be performed instantaneously. Thus, even when the allotment of computation resources is limited, the acoustic quality is not greatly impacted.
The updating of the spatial information may be periodically performed with each elapse of a preset time period or term, or may be performed when a preset condition is satisfied. Furthermore, the updating of the spatial information may be performed manually by the listener or a sound space manager, or may be performed by being triggered by a change to an external system.
For example, the spatial information may be updated when a controller is being operated by the listener and the position of the listener's avatar is instantaneously warped or time is instantaneously progressed or reversed. Alternatively, the spatial information may be updated when an operation to suddenly change the field environment is performed by the manager of the virtual space. In these cases, the thread for updating the spatial information managed by spatial information manager 1201 may be activated as one-time interrupt processing in addition to the periodic activation.
The role of the information update thread that performs update processing of the spatial information is, for example: processing to update, based on the position or orientation of the VR goggles worn by the listener, the position or orientation of the listener's avatar positioned in the virtual space; updating of the positions of objects that have moved within the virtual space; and the like. This is handled within a processing thread that operates at a relatively low frequency of around several tens of Hz. Processing that reflects the characteristics of direct sound may also be performed in such a low-frequency processing thread. This is because the characteristics of a direct sound vary less frequently than audio processing frames for audio output occur. Rather, by doing so, the computational load of such processing can be made relatively low, and the risk of pulsive noise can be avoided, since updating information at an unduly high frequency would generate the risk of pulsive noise occurring.
FIG. 4B is a block diagram illustrating another configuration example of a decoder. Specifically, FIG. 4B illustrates the configuration of decoder 1210, which is another example of decoder 1112 in FIG. 3B and FIG. 3D.
FIG. 4B is different from FIG. 4A in the respect that input data 1113 includes not encoded audio data, but an unencoded audio signal. Input data 1113 includes an audio signal and a bitstream including metadata.
Spatial information manager 1211 is the same as spatial information manager 1201 in FIG. 4A; therefore, description thereof has been omitted.
Renderer 1213 is the same as renderer 1203 in FIG. 4A; therefore, description thereof has been omitted.
Note that decoders 1112, 1200, and 1210 may be expressed as the acoustic processor that performs the acoustic processing. Furthermore, decoding devices 1110 and 1130 may be audio signal processing device 1001, or may be expressed as the acoustic processing device.
(Physical Configuration of Audio Signal Processing Device)
FIG. 5 is a diagram illustrating an example of a physical configuration of audio signal processing device 1001. Note that audio signal processing device 1001 in FIG. 5 may be decoding device 1110 in FIG. 3B or decoding device 1130 in FIG. 3D. A plurality of the constituent elements illustrated in FIG. 3B or FIG. 3D may be implemented by a plurality of the constituent elements illustrated in FIG. 5. Furthermore, a part of the configuration described here may be included in audio presentation device 1002.
Audio signal processing device 1001 in FIG. 5 includes processor 1402, memory 1404, communication interface (I/F) 1403, sensor 1405, and loudspeaker 1401.
Processor 1402 is, for example, a CPU, a digital signal processor (DSP), or a graphics processing unit (GPU). The acoustic processing or the decoding processing of the present disclosure may be performed by the CPU, the DSP, or the GPU executing a program stored in memory 1404. Furthermore, processor 1402 is, for example, a circuit that performs information processing. Processor 1402 may be a dedicated circuit that performs signal processing on audio signals, including the acoustic processing of the present disclosure.
Memory 1404 includes, for example, random access memory (RAM) or read-only memory (ROM). Memory 1404 may include, for example, magnetic storage media, exemplified by a hard disk, or semiconductor memory, exemplified by an SSD. Furthermore, memory 1404 may be an internal memory incorporated into the CPU or GPU. Moreover, spatial information managed by spatial information manager 1201 and 1211 and/or the like may be stored in memory 1404. Furthermore, threshold value data, described later, may be stored.
Communication I/F 1403 is, for example, a communication module that supports a communication method such as Bluetooth (registered trademark) or WiGig (registered trademark). Audio signal processing device 1001 communicates with other communication devices via communication I/F 1403, and obtains a bitstream to be decoded. The obtained bitstream is, for example, stored in memory 1404.
Communication I/F 1403 includes, for example, a signal processing circuit that supports the communication method, and an antenna. The communication method is not limited to Bluetooth (registered trademark) or WiGig (registered trademark), and may be Long Term Evolution (LTE), New Radio (NR), Wi-Fi (registered trademark), or the like.
The communication method is not limited to the wireless communication methods described above, and may be a wired communication method such as Ethernet (registered trademark), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI) (registered trademark), or the like.
Sensor 1405 performs sensing to estimate the position or orientation of the listener. Specifically, sensor 1405 estimates the position and/or orientation of the listener based on one or more detection results of one or more of the position, orientation, movement, velocity, angular velocity, acceleration, or the like of a part or all of the listener's body, and generates position/orientation information indicating the position and/or orientation of the listener.
Note that a device outside of audio signal processing device 1001 may include sensor 1405. The part of the body may be the listener's head or the like. The position/orientation information may be information indicating the position and/or orientation of the listener in real-world space, or may be information indicating the displacement of the position and/or orientation of the listener with respect to the position and/or orientation of the listener at a predetermined time point. Furthermore, the position/orientation information may be information indicating a position and/or orientation relative to three-dimensional sound reproduction system 1000 or an external device including sensor 1405.
Sensor 1405 may be, for example, an imaging device such as a camera or a distance measuring device such as a laser imaging detection and ranging (LIDAR) distance measuring device. Sensor 1405 may capture an image of the movement of the listener's head and detect the movement of the listener's head by processing the captured image. Furthermore, a device that performs position estimation using radio waves in any given frequency band such as millimeter waves may be used as sensor 1405.
Furthermore, audio signal processing device 1001 may obtain position information via communication I/F 1403 from an external device including sensor 1405. In this case, audio signal processing device 1001 need not include sensor 1405. Here, the external device refers to, for example, audio presentation device 1002 described in FIG. 2, or a stereoscopic image reproduction device worn on the listener's head. In this case, sensor 1405 is configured as a combination of various sensors, such as a gyro sensor and an acceleration sensor, for example.
As the speed of the movement of the listener's head, sensor 1405 may detect, for example, the angular speed of rotation about at least one of three mutually orthogonal axes in the sound space as the axis of rotation or the acceleration of displacement in at least one of the three axes as the direction of displacement.
As the amount of the movement of the listener's head, sensor 1405 may detect, for example, the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the axis of rotation or the amount of displacement in at least one of the three axes as the direction of displacement. Specifically, sensor 1405 detects 6DoF positions (x, y, z) and angles (yaw, pitch, roll) as the position of the listener. Sensor 1405 is configured as a combination of various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.
Note that sensor 1405 may implemented by, e.g., a camera or a Global Positioning System (GPS) receiver for detecting the position of the listener. Position information obtained by performing self-position estimation, by using LIDAR or the like as sensor 1405, may be used. For example, when three-dimensional sound reproduction system 1000 is implemented by a smartphone, sensor 1405 is included in the smartphone.
Furthermore, sensor 1405 may include a temperature sensor such as a thermocouple that detects the temperature of audio signal processing device 1001. Moreover, sensor 1405 may include, for example, a sensor that detects the remaining level of a battery included in audio signal processing device 1001 or a battery connected to audio signal processing device 1001.
Loudspeaker 1401 includes, for example, a diaphragm, a driving mechanism such as a magnet or a voice coil, and an amplifier, and presents the acoustic-processed audio signal as sound to the listener. Loudspeaker 1401 operates the driving mechanism according to the audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified via the amplifier, and vibrates the diaphragm by means of the driving mechanism. In this way, the diaphragm vibrating according to the audio signal generates sound waves, which propagate through the air and are transmitted to the listener's ears, allowing the listener to perceive the sound.
Note that although here, an example in which audio signal processing device 1001 includes loudspeaker 1401 and presents the acoustic-processed audio signal via loudspeaker 1401 was given, the means for providing the audio signal is not limited to this configuration.
For example, the acoustic-processed audio signal may be outputted to external audio presentation device 1002 connected via a communication module. The communication performed by the communication module may be wired or wireless. As another example, audio signal processing device 1001 may include a terminal that outputs an analog audio signal, and may present the audio signal from earphones or the like by connecting the earphone cable to the terminal.
In this case, audio presentation device 1002 may be headphones, earphones, a head-mounted display, neck speakers, wearable speakers, or the like, each worn on the listener's head or a part of the listener's body. Alternatively, audio presentation device 1002 may be surround speakers configured with a plurality of fixed speakers, or the like. Audio presentation device 1002 may reproduce the audio signal.
(Physical Configuration of Encoding Device)
FIG. 6 is a diagram illustrating an example of a physical configuration of encoding device 1500. Encoding device 1500 in FIG. 6 may be encoding device 1100 in FIG. 3A or encoding device 1120 in FIG. 3C, or a plurality of the constituent elements illustrated in FIG. 3A or FIG. 3C may be implemented by a plurality of the constituent elements illustrated in FIG. 6.
Encoding device 1500 in FIG. 6 includes processor 1501, memory 1503, and communication I/F 1502.
Processor 1501 is, for example, a CPU, a DSP, or a GPU. The encoding processing of the present disclosure may be performed by the CPU, the DSP, or the GPU executing a program stored in memory 1503. Furthermore, processor 1501 is, for example, a circuit that performs information processing. Processor 1501 may be a dedicated circuit that performs signal processing on audio signals, including the encoding processing of the present disclosure.
Memory 1503 includes, for example, RAM or ROM. Memory 1503 may include, for example, magnetic storage media, exemplified by a hard disk, or semiconductor memory, exemplified by an SSD. Furthermore, memory 1503 may be an internal memory incorporated into the CPU or GPU.
Communication I/F 1502 is, for example, a communication module that supports a communication method such as Bluetooth (registered trademark) or WiGig (registered trademark). For example, encoding device 1500 communicates with other communication devices via communication I/F 1502, and transmits an encoded bitstream.
Communication I/F 1502 includes, for example, a signal processing circuit that supports the communication method, and an antenna. The communication method is not limited to Bluetooth (registered trademark) or WiGig (registered trademark), and may be LTE, NR, Wi-Fi (registered trademark), or the like. The communication method is not limited to wireless communication methods. The communication method may be a wired communication method such as Ethernet (registered trademark), USB, HDMI (registered trademark), or the like.
The communication module includes, for example, a signal processing circuit that supports the communication method, and an antenna. In the above example, Bluetooth (registered trademark) and WiGig (registered trademark) were given as examples of the communication method, but a communication method such as Long Term Evolution (LTE), New Radio (NR), Wi-Fi (registered trademark), or the like may be supported. Furthermore, the communication I/F may be not the wireless communication methods described above, but a wired communication method such as Ethernet (registered trademark), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI) (registered trademark), or the like.
[Configuration of Renderer]
FIG. 7 is a block diagram illustrating a configuration example of renderer 1300. Specifically, FIG. 7 illustrates an example of the detailed configuration of renderer 1300, which corresponds to renderers 1203 and 1213 in FIG. 4A and FIG. 4B.
Renderer 1300 includes analyzer 1301, selector 1302, and reproducer 1303, and adds acoustic processing to sound data included in the input signal and outputs the sound data.
The input signal includes, for example, spatial information, sensor information, and sound data. The input signal may include a bitstream that includes sound data and metadata (control information), and in this case, the spatial information may be included in the metadata.
The spatial information is information related to the sound space (three-dimensional sound field) created by three-dimensional sound reproduction system 1000, and includes information about objects included in the sound space and information about the listener. The objects include sound source objects that emit sound and serve as sound sources, and non-sound-emitting objects that do not emit sound. The sound source objects may be expressed as simply sound sources.
The non-sound-emitting object serves as an obstacle object that reflects sound emitted by the sound source object, but a sound source object may also serve as an obstacle object that reflects sound emitted by another sound source object. The obstacle object may also be expressed as a reflection object.
Information assigned in common to both sound source objects and non-sound-emitting objects includes position information, geometry information, and the attenuation rate of sound volume when the object reflects sound.
The position information is represented by coordinate values of three axes, for example, the X-axis, the Y-axis, and the Z-axis of Euclidean space, but it does not necessarily have to be three-dimensional information. For example, the position information may be two-dimensional information represented by coordinate values of the two axes of the X-axis and the Y-axis. The position information of the object is defined by a representative position of the shape expressed by a mesh or voxel.
The geometry information may include information about the material of the surface.
The attenuation rate may be expressed as a real number greater than or equal to 0 and less than or equal to 1, or may be expressed as a negative decibel value. Since sound volume does not increase from reflection in real-world space, the attenuation rate is set to a negative decibel value. However, for example, to create an eerie atmosphere in a non-realistic space, an attenuation rate greater than or equal to 1, that is, a positive decibel value, may be intentionally set.
Furthermore, the attenuation rate may be set such that each frequency band included in a plurality of frequency bands has a different value, or values may be independently set for each frequency band. Furthermore, when the attenuation rate is set for each type of material of an object surface, the value of the corresponding attenuation rate may be used based on information about the surface material.
Furthermore, the spatial information may include, for example, information indicating whether the object belongs to an animate thing or information indicating whether the object is a mobile body. When the object is a mobile body, the position indicated by the position information may move over time. In this case, information on the changed position or the amount of change is transmitted to renderer 1300.
Information related to the sound source object includes, in addition to information assigned in common to both sound source objects and non-sound-emitting objects, sound data and information necessary for radiating the sound data into the sound space. The sound data is data representing sound perceived by the listener, and indicates information such as the frequency and intensity of the sound.
The sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, since the signal needs to be decoded at least before arriving at reproducer 1303, renderer 1300 may include a decoder (not illustrated). Alternatively, the signal may be decoded by audio data decoder 1202.
One item of sound data may be set for one sound source object, or a plurality of items of sound data may be set. Furthermore, identification information for identifying each item of sound data may be assigned to the sound data, or information related to the sound source object may include identification information for the sound data.
As information necessary for radiating sound data into the sound space, for example, information on a reference sound volume that is used as a standard in reproducing the sound data, information indicating a characteristic of sound data, information related to the position of the sound source object, information related to the orientation of the sound source object (in other words, information related to the directivity of the sound emitted by the sound source object), and the like may be included.
The information on the reference sound volume may be, for example, the root mean square value of the amplitude value of the sound data at the sound source position at the time of radiating the sound data into the sound space, and may be expressed as a floating-point decibel (dB) value.
For example, when the reference sound volume is 0 dB, the information on the reference sound volume may indicate that the sound is to be radiated into the sound space from the position indicated by the information related to the position of the sound source object, at the same sound volume as the signal level indicated by the sound data, without increase or decrease. Furthermore, when the reference sound volume is-6 dB, this may indicate that the sound is to be radiated into the sound space from the position indicated by the information related to the position of the sound source object at approximately half the sound volume of the signal level indicated by the sound data.
The information on the reference sound volume may be assigned to each item of sound data, or may be assigned collectively to a plurality of items of sound data.
The information indicating a characteristic of sound data is, for example, information on the sound volume of a sound source, and may be information indicating the chronological variation in the sound volume of a sound source.
For example, when the sound space is a virtual conference room and the sound source is a speaker, the sound volume transitions intermittently over short periods of time. In other words sound portions and silent portions occur alternately. Furthermore, when the sound space is a concert hall and the sound source is a performer, the sound volume is maintained over a certain duration of time. Moreover, when the sound space is a battlefield and the sound source is an explosive, the sound volume of the explosion sound becomes large for only an instant and then continues to be silent or in a quiet state thereafter.
In this way, the sound volume information on the sound source may include not only information on the magnitude of sound but also information on the transition of the sound magnitude. Such information may be used as the information indicating a characteristic of the sound data.
The information on the transition may be expressed by data indicating frequency characteristics in chronological order. The information on the transition may be expressed by data indicating the duration of a sound interval. The information on the transition may be expressed by data indicating the chronological order of durations of sound intervals and durations of silent intervals. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration for which the amplitude of the sound signal can be considered stationary (can be considered approximately constant) and the amplitude value of said signal during that duration.
The information on the transition may be expressed by data of a duration during which the frequency characteristics of the sound signal can be considered stationary. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration during which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during that duration. The information on the transition may be expressed in the format of, for example, data indicating the general shape of a spectrogram.
Furthermore, the sound volume that is used as the standard for the above-mentioned frequency characteristics may be used as the reference sound volume. The information on the reference sound volume and the information indicating a characteristic of the sound data may be used for calculation processing of the sound volume of direct sound or reflected sound to be perceived by the listener, and/or may be used for selection processing for selecting whether to make the listener perceive the sound. Other examples of and usage methods for the information indicating a characteristic of sound data will be described later.
It should be noted that the reflected sound according to the present embodiment is an example of indirect sound. Indirect sound may be reflected sound, diffracted sound, or the like. In the present embodiment, reflected sound, which is an example of indirect sound, is used for description, but the same processing is performed even if indirect sound other than reflected sound is used.
Information regarding the orientation of a sound source object (orientation information) is typically expressed in terms of yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted, and the orientation information of a sound source object may be expressed in terms of azimuth (yaw) and elevation (pitch). The orientation information of a sound source object may change over time, and when changed, the orientation information is transmitted to renderer 1300.
Information related to the listener is information regarding the position and orientation of the listener in the sound space. The information regarding the position (position information) is represented by the position on the X-, Y-, and Z-axes of Euclidean space, but need not necessarily be three-dimensional information and may be two-dimensional information. Information regarding the orientation of the listener (orientation information) is typically expressed in terms of yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted, and the listener orientation information may be expressed in terms of azimuth (yaw) and elevation (pitch).
The position information and orientation information regarding a listener may change over time, and when changed, the position information and orientation information are transmitted to renderer 1300.
The sensor information is information that includes, e.g., the rotation amount or displacement amount detected by sensor 1405 worn by the listener, and the position and orientation of the listener. The sensor information is transmitted to renderer 1300, and renderer 1300 updates the information on the position and orientation of the listener based on the sensor information. The sensor information may include position information obtained by performing self-localization estimation by a mobile terminal using GPS, a camera, or LIDAR, for example.
Furthermore, information obtained not from sensor 1405, but from an external source through a communication module, may also be detected as sensor information. Information indicating the temperature of audio signal processing device 1001, and information indicating the remaining level of the battery may be obtained from sensor 1405. Moreover, computational resources (CPU capability, memory resources, PC performance, and the like) of audio signal processing device 1001 or audio presentation device 1002 may be obtained in real time.
Analyzer 1301 analyzes an audio signal included in the input signal and spatial information received from spatial information managers 1201 and 1211 to calculate the information required for generating direct sounds and reflected sounds with reproducer 1303, and the information required for selecting whether to generate reflected sounds.
The information required for generating direct sounds and reflected sounds is, for example, for each of direct sounds and reflected sounds, values related to the path until arriving at the listening position, the time period taken until arrival, the sound volume at the arrival time, and the like. The values related to, e.g., the path until arriving at the listening position, the time period taken until arrival, and the sound volume at the arrival time are, for example, values representing the path until arriving at the listening position, the time period taken until arrival, and the sound volume at the arrival time, respectively.
The information required for selecting a reflected sound to be output is information indicating the relationship between the direct sound and the reflected sound, and is, for example, a value regarding a time difference between the direct sound and the reflected sound, a value regarding a sound volume ratio of the reflected sound to the direct sound at the listening position, and/or the like. The value related to the time difference between a direct sound and a reflected sound and the value related to the sound volume ratio between a direct sound and a reflected sound at the listening position are, for example, a value representing the time difference between the direct sound and the reflected sound and a value representing the sound volume ratio of the reflected sound to the direct sound at the listening position, respectively.
Note that it goes without saying that when the sound volume is expressed in units of decibels on a logarithmic scale (when the sound volume is expressed in the decibel domain), the sound volume ratio between the two signals is expressed as a decibel value difference. Specifically, the sound volume ratio between the two signals may be the difference when the amplitude value of each signal is expressed in the decibel domain. That value may be calculated based on, e.g., an energy value, a power value, or the like. Furthermore, this difference can be referred to as a difference in gain or simply a gain difference, in the decibel domain.
In other words, the sound volume ratio in the present disclosure is essentially the ratio between the amplitudes of signals; thus, the sound volume ratio may be expressed as a loudness ratio, a volume ratio, an amplitude ratio, a sound level ratio, a sound intensity ratio, a gain ratio, or the like. Furthermore, when the unit of sound volume is decibels, it goes without saying that the sound volume ratio in the present disclosure may be rephrased as the sound volume difference.
In the present disclosure, the “sound volume ratio” typically means the gain difference when the sound volume of each of two sounds is expressed in the unit of decibels, and in the examples of the embodiment, the threshold value data is also typically specified by the gain difference expressed in the decibel domain. However, the sound volume ratio is not limited to the gain difference in the decibel domain. When a sound volume ratio that is not expressed by the decibel domain is used, threshold value data specified in the decibel domain may be used by converting the threshold value data into the unit of the sound volume ratio calculated. Alternatively, threshold value data specified beforehand in each unit may be stored in the memory.
In other words, for example, even if a ratio between energy values, power value, or the like is used instead of the sound volume ratio, it is obvious that the algorithm in the present disclosure can be applied to solve the problem of the present disclosure.
The time difference between the arrival of a direct sound and the arrival of a reflected sound is, for example, the time difference between a direct sound arrival time period (arrival time) and a reflected sound arrival time period (arrival time). It should be noted that for simplicity, the time difference between the arrival of a direct sound and the arrival of a reflected sound may be referred to as the time difference between the direct sound and the reflected sound. The time difference between a direct sound and a reflected sound may be the time difference between the times at which each of the direct sound and the reflected sound arrive at the listening position, the difference in the time periods taken until each of the direct sound and the reflected sound arrive at the listening position, or the time difference between the time when emission of the direct sound ends and the time when the reflected sound arrives at the listening position. The methods for calculating these values will be described later.
Selector 1302 selects whether reproducer 1303 is to generate a reflected sound by using information calculated by analyzer 1301 and the threshold value data. To put it differently, selector 1302 determines whether to select a reflected sound as a reflected sound to be generated. To put it still differently, selector 1302 selects which reflected sounds reproducer 1303 is to generate, from a plurality of reflected sounds.
The threshold value data is, for example, a graph having a horizontal axis that indicates the time difference between a direct sound and reflected sounds and a vertical axis that indicates the sound volume ratios of reflected sounds to a direct sound, and is expressed as a boundary (threshold value) that demarcates whether each reflected sound is perceived. For example, the threshold value data may be expressed as an approximation formula that includes the time difference between a direct sound and a reflected sound as a variable, or may be expressed as an arrangement that includes values of time differences between direct sounds and reflected sounds as an index, and corresponding threshold values.
Selector 1302 selects the generation of a reflected sound when, for example, at the time difference between the arrival time of a direct sound and the arrival time of a reflected sound, the sound volume ratio of the arrival time sound volume of the reflected sound to the arrival time sound volume of the direct sound is a value that is larger than a threshold value set with reference to threshold value data. It should be noted that the arrival time sound volume means the sound volume when the sound arrives at the listening position.
To put it differently, the time difference between the arrival time of a direct sound and the arrival time of a reflected sound is the difference in the amount of time taken for the direct sound and the reflected sound to arrive at the listening position. Furthermore, the time difference between the time point at which emission of the direct sound stops and the time point at which the reflected sound arrives at the listening position may be used as the time difference between the direct sound and the reflected sound. In this case, threshold value data that is different from the threshold value data determined by using, as a reference, the time difference between the direct sound arrival time and the reflected sound arrive time may be used, or common threshold value data may be used.
The threshold value data may be obtained from memory 1404 of audio signal processing device 1001, or may be obtained from an external storage device via a communication module. The threshold value data storage method and the threshold value setting method will be described later.
Reproducer 1303 synthesizes the audio signals of direct sounds and the audio signals of reflected sounds selected for generation by selector 1302.
Specifically, reproducer 1303 processes the inputted audio signals to generate direct sounds, based on information on the direct sound arrival time and the direct sound arrival time sound volume calculated by analyzer 1301. Furthermore, reproducer 1303 processes the inputted audio signals to generate reflected sounds, based on information on the reflected sound arrival time and the reflected sound arrival time sound volume pertaining to the reflected sounds selected by selector 1302. Then, reproducer 1303 synthesizes and outputs the direct sounds and reflected sounds that were generated.
[Operation Example of Renderer]
FIG. 8 is a flowchart illustrating an operation example of audio signal processing device 1001. FIG. 8 illustrates the processing performed mainly by renderer 1300 of audio signal processing device 1001.
In the analysis processing of the input signal (S101 in FIG. 8), analyzer 1301 analyzes the input signal inputted into audio signal processing device 1001 to detect direct sounds and reflected sounds that may be generated in the sound space. The reflected sounds detected here are candidates for the reflected sounds to be selected by selector 1302 as the reflected sounds to be ultimately generated by reproducer 1303. Furthermore, analyzer 1301 analyzes the input signal to calculate information necessary for generating direct sound and reflected sound, and information necessary for selecting the reflected sounds to be generated.
First, the characteristics of each of the direct sound and the reflected sound are calculated. Specifically, the arrival time period and the arrival time sound volume when each of the direct sound and the reflected sound arrive at the listener are calculated. When a plurality of objects are present in the sound space as reflection objects, reflected sound characteristics with respect to each of the plurality of objects are calculated.
The direct sound arrival time period (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is a path that connects position information(S) (xs, ys, zs) of a sound source object with position information A1 (xa, ya, za) of the listener. The direct sound arrival time period (td) is a value obtained by dividing the length of the path that connects position information(S) (xs, ys, zs) with position information A1 (xa, ya, za), by the speed of sound (approximately 340 m/sec).
For example, the path length (X) is determined by the expression ((xs−xa){circumflex over ( )}2+(ys−ya){circumflex over ( )}2+(zs−za){circumflex over ( )}2){circumflex over ( )}0.5. The sound volume attenuates in inverse proportion to the distance. Thus, when the sound volume at position information S (xs, ys, zs) of a sound source object is denoted by N and the unit distance is denoted by U, the direct sound arrival time sound volume (Id) is determined by the expression Id=N*U/X.
Sound volume N at the sound source position may be the reference sound volume described above.
The reflected sound arrival time period (tr) is calculated based on the reflected sound arrival path (pr). The reflected sound arrival path (pr) is a path that connects the position of the sound image of a reflected sound with position information A1 (xa, ya, za).
Note that the position of the sound image of the reflected sound may be derived by using, for example, a “mirror image method” or a “ray tracing method”, or by using any other method for deriving sound image positions. The mirror image method is a method that simulates a sound image by assuming that a reflected wave on the wall in a room has a mirror image in a position symmetrical to the sound source relative to the wall, and that sound waves are emitted from the position of that mirror image. The ray tracing method is a method that simulates, for example, an image (sound image) observed at a certain point by tracing waves that are transmitted in a linear manner, such as light rays or sound rays.
FIG. 9 is a diagram illustrating a comparatively distant positional relationship between a listener and an obstacle object. FIG. 10 is a diagram illustrating a comparatively close positional relationship between a listener and an obstacle object. In other words, each of FIG. 9 and FIG. 10 illustrate an example in which the sound image of a reflected sound is formed in a position symmetrical to the sound source position, with a wall interposed therebetween. Based on such a relationship, by determining the position of the sound image of the reflected sound on the x-, y-, and z-axes, the reflected sound arrival time period can be determined in the same manner as the method for calculating the direct sound arrival time period.
The reflected sound arrival time period (tr) is a value obtained by dividing the length of the path that connects the position of the sound image of a reflected sound with position information A1 (xa, ya, za), by the speed of sound (approximately 340 m/sec). The sound volume attenuates in inverse proportion to the distance. Thus, when the sound volume at the sound source position is denoted by N, the unit distance is denoted by U, and the attenuation rate of the sound volume at the reflection is denoted by G, the reflected sound arrival time sound volume (Ir) is determined by the expression Ir=N*G*U/Y.
As described above, attenuation rate G may be expressed as a real number greater than or equal to 0 and less than or equal to 1, or may be expressed as a negative decibel value. In this case, the sound volume of the signal as a whole attenuates by the amount of G. Furthermore, the attenuation rate may be set for each frequency band included in a plurality of frequency bands. In this case, analyzer 1301 multiplies each frequency component of the signal by the specified attenuation rate. Furthermore, in order to reduce the amount of computation, analyzer 1301 may, by using, as an overall attenuation rate, a representative value, an average value, or the like of a plurality of attenuation rates of a plurality of frequency bands, cause the sound volume of the signal as a whole to attenuate by that amount.
Next, analyzer 1301 calculates the sound volume ratio (L), which is the ratio of the reflected sound arrival time sound volume (Ir) to the direct sound arrival time sound volume (Id), and the time difference (T) between the direct sound and the reflected sound, each of the sound volume ratio (L) and the time difference (T) being required for selection of the reflected sound to be generated.
The sound volume ratio (L), which is the ratio of the above-described Ir to the direct sound arrival time sound volume (Id), is, for example, a value obtained by dividing the reflected sound arrival time sound volume (Ir) by the direct sound arrival time sound volume (Id), and is determined by: L=(N*G*U/Y)/(N*U/X)=G*X/Y. Since the value to be determined is a sound volume ratio, the values of N and U may be any predetermined values.
The time difference (T) between a direct sound and a reflected sound may be, for example, the time difference between the time periods each of the direct sound and the reflected sound take to arrive at the listening position. For example, the time difference (T) between the time periods taken for each of a direct sound and a reflected sound to arrive at the listening position is determined by T=tr−td.
Furthermore, the time difference (T) may be the difference between the times at which each of a direct sound and a reflected sound arrive at the listening position. Moreover, the time difference (T) may be the time difference between the time at which the emission of the direct sound ends and the time at which the reflected sound arrives at the listening position. In other words, the time difference (T) may be the time difference, at the listening position, between the time at which the direct sound ends and the time at which the reflected sound begins.
Next, in reflected sound selection processing (S102 in FIG. 8), selector 1302 selects whether reproducer 1303 is to generate a reflected sound calculated by analyzer 1301. To put it differently, selector 1302 determines whether to select a reflected sound as a reflected sound to be generated. When there are a plurality of reflected sounds, selector 1302 selects whether to generate each reflected sound. As the result of selecting whether to generate each reflected sound, selector 1302 may select one or more reflected sounds to be generated from the plurality of reflected sounds, or may select one reflected sound to be generated.
Note that selector 1302 may select reflected sounds to which other processing is to be applied, not limited to generation processing. For example, selector 1302 may select reflected sounds to which binaural processing is to be applied. Furthermore, selector 1302 fundamentally selects only the one or more reflected sounds that are to be processed. However, selector 1302 may select only one or more reflected sounds that are not to be processed. Processing may then be applied to the one or more reflected sounds that were not selected.
For example, the selection of reflected sounds may be performed based on the sound volume ratio (L) and the time difference (T) calculated by analyzer 1301. Due to the selection processing being performed based on the time difference (T) between direct sounds and reflected sounds, it is possible to more appropriately select reflected sounds that have a large degree of influence on the listener's perception, in comparison to when performing the selection processing based only on the sound volume difference between direct sounds and reflected sounds.
Specifically, the selection of whether to generate a reflected sound is performed by comparing, to a preset threshold value, the sound volume ratio of a reflected sound to a direct sound, the sound volume ratio corresponding to the time difference between the direct sound and the reflected sound. The threshold value is set with reference to the threshold value data. The threshold value data is an indicator indicating the boundary that demarcates whether a reflected sound corresponding to a direct sound is perceived by the listener, and is defined as the ratio between the direct sound arrival time sound volume (Id) and the reflected sound arrival time sound volume (Ir).
Note that the threshold value corresponds to a value expressed by, e.g., a numerical value determined based at the time difference (T). The threshold value data corresponds to the relationship between the time difference (T) and a threshold value, and corresponds to table data or a relational expression used for identifying or calculating the threshold value at the time difference (T). The format and type of the threshold value data is not limited to table data or a relational expression.
FIG. 11 is a diagram illustrating relationships between time differences between direct sounds and reflected sounds, and threshold values. For example, threshold value data of predetermined sound volume ratios may be referenced for each value of the time difference between a direct sound and a reflected sound, as illustrated in FIG. 11. Alternatively, threshold value data obtained by, e.g., interpolating or extrapolating from the threshold value data illustrated in FIG. 11 may be referenced.
Furthermore, the threshold value of the sound volume ratio at the time difference (T) calculated by analyzer 1301 is identified from the threshold value data. Moreover, selector 1302 determines whether to select a reflected sound as a reflected sound to be generated based on whether the sound volume ratio (L) of the reflected sound to the direct sound calculated by analyzer 1301 exceeds the threshold value.
Due to performing the selection processing by using the threshold value data of the sound volume ratio that is predetermined for each value of the time difference between a direct sound and a reflected sound, selection processing that considers post-masking or the precedence effect can be achieved. The type, format, storage method, setting method, and the like of the threshold value data will be described in detail later.
Next, in the generation processing of direct sounds and reflected sounds (S103 in FIG. 8), reproducer 1303 generates and synthesizes the audio signals for direct sounds and the audio signals for reflected sounds that have been selected by selector 1302 as reflected sounds to be generated.
The audio signals for direct sounds are generated by applying the direct sound arrival time period (td) and the direct sound arrival time sound volume (Id) calculated by analyzer 1301 to the sound data for the sound source objects included in the input signal. Specifically, processing in which the sound data is delayed by the amount of the direct sound arrival time period (td) and multiplied by the direct sound arrival time sound volume (Id) is performed. The processing to delay the sound data is processing in which the position of the sound data is moved forward or backward on the time axis. Processing in which the sound data is delayed may be applied without causing the sound quality to deteriorate, such as was disclosed in PTL 2.
The audio signals for reflected sounds are, similarly to the direct sounds, generated by applying the reflected sound arrival time period (tr) and the reflected sound arrival time sound volume (Ir) calculated by analyzer 1301 to the sound data for the sound source objects.
However, the reflected sound arrival time sound volume (Ir) in the generation of reflected sounds differs from the arrival time sound volume of the direct sounds in that the arrival time sound volume of the reflected sounds is a value to which attenuation rate G of the sound volume in the reflection has been applied. G may be an attenuation rate that is applied globally to all frequency bands. Alternatively, in order to reflect the biases of frequency components generated by reflection, the reflectance may be defined for each predetermined frequency band. In this case, the processing to apply the reflected sound arrival time sound volume (Ir) may be performed as frequency equalizer processing, which is processing that involves multiplying each band by the attenuation rate.
In the above example, for each of the direct sounds and the reflected sound candidates, the path length when arriving at the listener is calculated. Furthermore, the arrival time period and the arrival time sound volume are calculated based on each path length. The selection processing of the reflected sound candidates is then performed based on the time differences and the sound volume ratios of these.
Note that as a different example, the selection processing may be performed based on the path lengths when each of the direct sound and the reflected sound arrive at the listener, and the calculation of the arrival time period and the arrival time sound volume of each of the direct sound and the reflected sound and the calculation of the time difference and the sound volume ratio may be omitted. In this case, threshold values according to path length differences may be determined beforehand with respect to path length ratios. Then, selection processing may be performed based on whether the path length ratio calculated is greater than or equal to the threshold value according to the path length difference calculated. This makes it possible to perform selection processing based on path length differences that correspond to time differences, while reducing the amount of computation.
Furthermore, a parameter that indicates sound propagation speed or a parameter that has an impact on the sound propagation speed parameter may be used in addition to the path length difference.
(Details of Selection Processing)
The selection processing that determines whether reflected sounds are generated will be explained in detail.
The selection of a reflected sound is performed by comparing, with the sound volume ratio (L) calculated by analyzer 1301, the threshold value determined for the sound volume ratio, which is the ratio of the reflected sound arrival time sound volume to the direct sound arrival time sound volume, at the time difference (T) between the direct sound and the reflected sound. For example, of threshold values of sound volume ratios that were determined beforehand for each value of a time difference between a direct sound and a reflected sound, the threshold value of the sound volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by analyzer 1301 is referenced. Then, determination of whether the reflected sound is selected as a reflected sound to be generated is made based on whether the sound volume ratio (L) calculated by analyzer 1301 exceeds the threshold value.
The time difference (T) may be any of, for example, the difference in the times at which each of a direct sound and a reflected sound arrive at the listening position, the time difference between the time periods taken when each of a direct sound and a reflected sound arrive at the listening position, or the time difference between the time point when emission of a direct sound stops and the time point when a reflected sound arrives at the listening position. Here, the direct sound end time may be determined by adding the duration of a direct sound to the arrival time of the direct sound.
The threshold value data may be determined based on the minimum time difference at which the perception of the listener is able to detect the divergence of two sounds due to an action of the auditory nerve or a cognitive effect in the brain, and more specifically due to the precedence effect, described later, the temporal masking phenomenon, described later, or a combination of both. Specific numerical values may be derived from research results into the temporal masking effect, the precedence effect, the echo detection limit, etc. that are already known, or may be determined by an auditory test performed with the premise of application in the virtual space.
FIG. 12A, FIG. 12B, and FIG. 12C are diagrams illustrating examples of threshold value data setting methods. As illustrated in FIG. 12A, FIG. 12B, and FIG. 12C, the threshold value data represents the boundaries (threshold values) determining whether reflected sound is perceived or not perceived, in a graph having a horizontal axis that indicates the time difference between direct sound and reflected sound and a vertical axis that indicates the sound volume ratio of the reflected sound to the direct sound.
The threshold value data may be expressed by an approximation formula that includes the time difference between direct sound and reflected sound as a variable. Furthermore, as illustrated in FIG. 11, the threshold value data may be stored in the domain of memory 1404 as an arrangement of an index of time differences between direct sounds and reflected sounds, and threshold values corresponding to the index.
Note that when the height of a line parallel to the horizontal axis in Example 4 in FIG. 12C (the minimum audible limit) is used as the threshold value, what is compared to the threshold value is not the sound volume ratio (L) between a direct sound and a reflected sound, but the sound volume of the reflected sound itself. The reason for this is that the threshold value indicates the sound volume of the boundary demarcating whether a sound is perceivable to the listener, and is a threshold value for determining a sound having a lower sound volume than the threshold value to be a sound that is not to be reproduced. In other words, the threshold value corresponding to the minimum audible limit is not a threshold value with respect to the ratio of the sound volume of a reflected sound to the sound volume of a direct sound.
When the minimum audible limit is used as the threshold value, the threshold value is constant regardless of the time difference (T); thus, the time difference (T) need not be calculated.
Note that when a plurality of reflected sounds are generated in the analysis processing (S101 in FIG. 8), the selection processing may be performed on all of the reflected sounds, or the selection processing may be performed on only the reflected sounds having high evaluation values based on the evaluation values derived for each reflected sound by means of a preset evaluation method. Here, the evaluation value of a reflected sound corresponds to the sensory level of importance of the reflected sound. Note that the evaluation value being high corresponds to the evaluation value being large, and these expressions may be used interchangeably.
Selector 1302 may calculate an evaluation value for each reflected sound by an evaluation method set beforehand based on, for example, the sound volume of the sound source, the visual properties of the sound source, the positionality of the sound source, the visual properties of the reflection object (the obstacle object), the geometrical relationship between the direct sound and the reflected sound, and/or the like.
Specifically, the evaluation value may become higher as the sound volume of the sound source is greater. Furthermore, in order to cause visual positioning and acoustic positioning to match each other, the evaluation value may be high when a sound source object or a reflection object (obstacle object) is visible from the listener, or when the positionality of a sound source object is high.
Moreover, the size of the arrival angle formed by a direct sound and a reflected sound and the difference between the arrival time periods of a direct sound and a reflected sound greatly affect the grasping of the space. Thus, the evaluation value may be high when the size of the angle formed by the arrival of a direct sound and the arrival of a reflected sound is large, or when the difference between the arrival time periods of a direct sound and a reflected sound is large.
The sound volume information on the sound source may indicate the reference sound volume determined for each content, a temporal transition of sound volume, or both.
For example, when the virtual space is a virtual conference room and the direct sound is a speaking voice, the sound volume transitions intermittently over short periods of time. In other words, sound portions and silent portions occur alternately. Furthermore, when the virtual space is a concert hall and the direct sound is the performance of a musical piece, the sound volume is maintained over a certain duration of time. Moreover, when the virtual space is a battlefield and the direct sound is an explosion sound, the sound volume becomes large for only an instant and then continues to be silent or in a quiet state thereafter.
In this way, the sound volume information on the sound source may include not only information on the reference sound volume corresponding to the volume setting when the sound is radiated into the virtual space, but also information on the transition of the sound magnitude.
The information on the transition may be expressed by data indicating frequency characteristics in chronological order. The information on the transition may be expressed by data indicating the duration of a sound interval. The information on the transition may be expressed by data indicating the chronological order of durations of sound intervals and durations of silent intervals. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration for which the amplitude of the sound signal can be considered stationary (can be considered approximately constant) and the amplitude value of said signal during that duration.
The information on the transition may be expressed by data of a duration during which the frequency characteristics of the sound signal can be considered stationary. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration for which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during that duration.
Furthermore, efforts to use temporal transitions in the frequency characteristics of signals for the acoustic processing of virtual spaces have been conventionally widely performed (PTL 1, etc.). When considering such conventional techniques, it goes without saying that the above-described sets may be sets that include durations of time in which the frequency characteristics are constant and those frequency characteristics.
The geometrical relationships may be positional relationships between a sound source, the listener, and a reflection object in a virtual space. The path lengths of each of the direct sound and the reflected sound arriving can be geometrically calculated based on these relationships. Therefore, using the relationship that sound volume is inversely proportional to distance, it is possible to calculate the reference sound volume of a reflected sound with respect to the reference sound volume of a direct sound.
A reflection coefficient of the reflection object may be used for calculation of the reference sound volume of a reflected sound. Furthermore, a generally used typical value may be used as the reflection coefficient. On the other hand, when special conditions are present, such as the reflection object being covered by a sound-absorbing material or the like, a specially assigned reflection coefficient may be used as the reflection coefficient of the reflection object.
Each reflected sound may be evaluated based on the sound volume of the reflected sound. The sound volume of the reflected sound may, as described above, be determined from the geometrical relationship between the direct sound and the reflected sound, and an indicator assigned to the reflection object. The reflected sound may be evaluated by comparing that sound volume with a predetermined threshold value.
Furthermore, information that indicates a temporal transition in the sound volume of the sound source may be reflected in the evaluation. For example, in a case in which the information that indicates a temporal transition in the sound volume of the sound source indicates the duration of a sound interval, the reflected sound evaluation value may be maintained as-is when the time is within the sound interval. On the other hand, when the time is outside of the sound interval, processing to reduce or make zero the evaluation value of the reflected sound may be performed, even if the reference sound volume of the reflected sound exceeds the threshold value.
Alternatively, the information that indicates a temporal transition in the sound volume of the sound source may be data that lists, in chronological order, a plurality of sets of a duration during which it is considered that the amplitude of the sound signal is largely constant and amplitude values of the signal during that period. In that case, processing may be performed such that reflected sounds are evaluated by changing the reference sound volume of the reflected sounds in coordination with changes in the amplitude values in the data.
Furthermore, both the information on the reference sound volume and the information on the temporally transitioning sound volume may be used as the information indicating the sound volume of a direct sound. For example, an evaluation value can be calculated based on the information on the reference sound volume, and then the evaluation value can be corrected by using the information on the transitioning sound volume.
In the evaluation of reflected sounds, all of the methods described above may be performed, or only some of the methods may be performed. For example, reflected sounds may be evaluated using a plurality of evaluation methods, or reflected sounds may be evaluated using one evaluation method.
When reflected sounds are evaluated using a plurality of evaluation methods, whether to select each reflected sound may be determined based on an evaluation value comprehensively determined by the plurality of evaluation methods, or may be determined based on the evaluation value of each of the plurality of evaluation methods.
When whether to select each reflected sound is determined based on each of the plurality of evaluation methods, audio signal processing device 1001 may select a sound when all of the plurality of evaluation results based on the plurality of evaluation methods indicate selecting the sound. Alternatively, audio signal processing device 1001 may select a sound when any one of the plurality of the evaluation results based on the plurality of evaluation methods indicates selecting the sound.
Furthermore, for example, an order of priority may be assigned to the first to third evaluation methods. When it is determined by the first evaluation method that a sound is not to be selected, audio signal processing device 1001 may then make a final determination that the sound is not to be selected, without depending on the determination results in the second and third evaluation methods. Moreover, when, although it has been determined by one of the second and third evaluation methods that a sound is not to be selected, it has been determined by the other of the second and third evaluation methods that the sound is to be selected, audio signal processing device 1001 may make a final determination that the sound is to be selected.
Furthermore, the selection processing and the evaluation processing may be performed independently of each other, or only one of these may be performed. Moreover, the evaluation processing may be performed only on reflected sounds that have been determined to be selected in the selection processing, and whether to select each reflected sound may be redetermined in the evaluation processing. Alternatively, the evaluation processing may be performed only on reflected sounds that have been determined to not be selected in the selection processing, and whether to select each reflected sound may be redetermined in the evaluation processing.
The selection processing described above can be interpreted as processing in which a reflected sound is selected in accordance with a characteristic of a direct sound. For example, in processing in which a reflected sound is selected in accordance with a characteristic of a direct sound, the threshold value used in selection of the reflected sound is set or adjusted in accordance with a characteristic of the direct sound. Alternatively, the evaluation value used in the selection of a reflected sound may be calculated based on one or more of, for example, the sound volume of the sound source, the visual properties of the sound source, the positionality of the sound source, the visual properties of the reflection object (the obstacle object), the geometrical relationship between the direct sound and the reflected sound, and/or the like.
Furthermore, the processing in which a reflected sound is selected based on a characteristic of a direct sound is not limited to processing in which the threshold value is set or adjusted in accordance with a characteristic of the direct sound and processing in which the evaluation value used for selection of the reflected sound to be processed is calculated, and other processes may be performed. Furthermore, even when performing the processing in which the threshold value is set or adjusted in accordance with a characteristic of the direct sound or the processing in which the evaluation value used in selection of the reflected sounds to be processed is calculated, the processing may be partially changed, or new processing may be added.
Note that setting the threshold value may include adjusting the threshold value, changing the threshold value, and the like.
[Threshold Value Setting Method]
The threshold value data used in the selection processing may be set with reference to the value of an echo detection limit based on a known precedence effect or a masking threshold value based on the post-masking effect.
The precedence effect is a phenomenon in which, when sounds are heard from two locations, it is perceived that the sound source is present at the location from which the first sound was heard. If two short sounds fuse together to be heard as one sound, the position (localization position) from which the overall sound is heard is, for the most part, determined by the position of the first sound. The echo detection limit is a phenomenon that occurs due to the precedence effect, and is the minimum time difference at which the listener's perception detects the divergence of two sounds.
In Example 2 of FIG. 12C, the horizontal axis corresponds to the arrival time period of reflected sound (echo), and specifically corresponds to the delay time period from the arrival time of direct sound to the arrival time of reflected sound. The vertical axis corresponds to the sound volume ratio of detectable reflected sound to direct sound, and specifically corresponds to the threshold value that determines whether reflected sound that has arrived with a delay time period is detectable.
FIG. 13 is a diagram illustrating an example of a threshold value setting method. The horizontal axis in FIG. 13 corresponds to the arrival time period of reflected sound, and specifically corresponds to the time differences (T) between direct sound and reflected sound. The vertical axis in FIG. 13 corresponds to the sound volume of reflected sound. Specifically, the vertical axis in FIG. 13 may correspond to the sound volume (sound volume ratios) of reflected sound determined in relation to direct sound, or may correspond to the sound volume of reflected sound determined absolutely without depending on the sound volume of the direct sound.
For example, when, as illustrated in FIG. 9, the listener and an obstacle object are comparatively far from each other, the arrival time period of the reflected sound becomes longer, and, as illustrated in C in FIG. 13, the threshold value is set to be low. As a result, in the case of FIG. 9, the reflected sound is generated. On the other hand, when, as illustrated in FIG. 10, the listener and the obstacle object are comparatively close to each other, the arrival time period of the reflected sound is shorter than that in the case of FIG. 9, and as illustrated in B in FIG. 13, the threshold value is set to be high. As a result, in the case of FIG. 10, the reflected sound is not generated.
Furthermore, the threshold value data may be stored in memory 1404, obtained from memory 1404 at the time of the selection processing, and used in the selection processing.
FIG. 14 is a flowchart illustrating an example of selection processing. First, selector 1302 specifies a reflected sound detected by analyzer 1301 (S201). Selector 1302 then detects the sound volume ratio (L) of the reflected sound to the direct sound, and the time difference (T) between the direct sound and the reflected sound (S202 and S203).
The time difference (T) may be any of, for example, the time difference between the time periods each of the direct sound and the reflected sound take to arrive at the listening position, the time difference between the direct sound arrival time and the reflected sound arrival time, and the time difference between the time when emission of the direct sound ends and the time when the reflected sound arrives at the listening position. Here, an example will be described based on the time difference between the direct sound arrival time and the reflected sound arrival time.
Specifically, based on: the position information on the sound source object and the listener; and the position information and geometry information on the obstacle object, selector 1302 calculates the difference between the length of the path of the direct sound and the length of the path of the reflected sound. By dividing the difference between the lengths by the speed of sound, selector 1302 then detects the time difference (T) between the time when the direct sound arrives at the listener's position and the time when the reflected sound arrives at the listener's position.
The sound volume when arriving at the listener attenuates, with respect to the sound volume of the sound source, in proportion to the distance to the listener (in inverse proportion to the distance). Therefore, the sound volume of the direct sound is obtained by dividing the sound volume of the sound source by the length of the path of the direct sound. The sound volume of the reflected sound is obtained by dividing the sound volume of the sound source by the length of the path of the reflected sound, and then further multiplying by the attenuation rate assigned to the virtual obstacle object. Selector 1302 detects the sound volume ratio by calculating the ratio between these sound volumes.
Furthermore, using the threshold value data, selector 1302 identifies the threshold value corresponding to the time difference (T) (S204). Selector 1302 then determines whether the sound volume ratio (L) detected is greater than or equal to the threshold value (S205).
When the sound volume ratio (L) is greater than or equal to the threshold value (“Yes” in S205), selector 1302 selects the reflected sound as a reflected sound to be generated (S206). When the sound volume ratio (L) is less than the threshold value (“No” in S205), selector 1302 skips selecting the reflected sound as a reflected sound to be generated (S207). That is, in this case, selector 1302 determines the reflected sound to be a reflected sound that is not to be generated, i.e., a reflected sound to be culled.
Subsequently, selector 1302 determines whether there are any unspecified reflected sounds (S208). If there are unspecified reflected sounds (“Yes” in S208), selector 1302 repeats the above-described processing (S201 to S207). If there are no unspecified reflected sounds (“No” in S208), selector 1302 ends the processing.
This selection processing may be performed on all of the reflected sounds generated in the analysis processing, or may be performed on only the reflected sounds for which the above-described evaluation value is high.
[Details of Threshold Value Storing Method]
The threshold value data according to the present embodiment is stored in memory 1404 of audio signal processing device 1001. The format and type of the threshold value data to be stored may be any format and any type. When threshold values having a plurality of formats and a plurality of types are stored, in the selection processing, the format and the type of the threshold values to be used in the selection processing of the reflected sounds may be decided. The method for determining which items of threshold value data to use in the selection processing will be described later.
Furthermore, a plurality of formats and a plurality of types of threshold value data may be stored in combination. The combined threshold value data may be read from spatial information managers 1201 and 1211 to set the threshold values to be used in the selection processing. Note that the threshold value data to be stored in memory 1404 may be stored in spatial information managers 1201 and 1211.
For example, the threshold value data may be stored as threshold values at each time difference, so as to plot a line between the threshold values as illustrated in [Example 1] and [Example 2] of FIG. 12C.
Furthermore, the threshold value data may be stored as table data in which, as illustrated in FIG. 11, the threshold values and the time differences (T) are associated with each other. In other words, the threshold value data may be stored as table data that includes the time differences (T) as an index. Naturally, the threshold values illustrated in FIG. 11 are examples, and the threshold values are not limited to the examples in FIG. 11. Furthermore, the threshold values may be approximated by functions that include the time differences (T) as variables, and coefficients of the functions may be stored, without storing the threshold values themselves. Moreover, a plurality of approximation expressions may be combined and stored.
For example, the threshold value data may be expressed by a formula such as that shown below, where timeDiff denotes the time difference (T), and gainThresh denotes the threshold value.
The threshold value is defined only over the time range in which the precedence effect is considered to occur. When the time difference is outside of that time range (in the above formula, a value of 1 ms or less or 40 ms or greater), the determination may be performed not by gainThresh, but solely by a threshold value representing the minimum sound volume for reproduction in the virtual space, described later.
Experiments by the present inventors have clarified that in the time range in which the precedence effect is considered to occur, the threshold value may be approximated by a function that is convex upward. The above-described formula is an example of an approximation formula generated based on these experiments.
Information on a relational expression that indicates the relationship between time differences (T) and threshold values may be stored in memory 1404. In other words, an expression that includes the time difference (T) as a variable may be stored. The threshold values of the time differences (T) may be approximated by a straight line or a curved line, and a parameter that indicates the geometrical shape of the straight line or the curved line may be stored. For example, when the geometrical shape is a straight line, the start point and the slope for expressing the straight line may be stored.
Furthermore, the threshold value data may be stored having the type and format thereof defined for each characteristic of direct sound. Moreover, parameters for adjusting threshold values based on a characteristic of the direct sound and using the threshold values in the selection processing may be stored. Processing to adjust threshold values in accordance with a characteristic of the direct sound and use the threshold values in the selection processing is described later, as a variation of the threshold value setting method.
As an example in which a plurality of types of threshold value data are stored in combination, as illustrated in [Example 3] in FIG. 12C, for each time difference (T), the larger value of the masking threshold value and the echo detection limit threshold value may be stored. As illustrated in [Example 4] in FIG. 12C, for each time difference (T), the larger value of the minimum sound volume for reproduction in a virtual space and the echo detection limit threshold value may be stored.
The combination of the plurality of types of the threshold value data is not limited to these. For example, in a plurality of items of threshold value data, information on the maximum value may be stored for each time difference (T).
Furthermore, in the above description, the information on threshold values has time period items as a one-dimensional index. The information on threshold values may have a two-dimensional or three-dimensional index that further includes variables related to the direction of arrival.
FIG. 15 is a diagram illustrating relationships between directions of direct sounds, directions of reflected sounds, time differences, and threshold values. For example, as illustrated in FIG. 15, threshold values pre-calculated in accordance with the relationship between the direct sound direction (θ), the reflected sound direction (γ), the time difference (T), and the sound volume ratio (L) may be stored.
The direct sound direction (θ) corresponds to the angle, with respect to the listener, of the direction of arrival of a direct sound. The reflected sound direction (γ) corresponds to the angle, with respect to the listener, of the direction of arrival of a reflected sound. Here, the direction in which the listener is facing is defined as 0 degrees. The time difference (T) corresponds to the difference, until reaching the listening position, of the direct sound arrival time period and the reflected sound arrival time period. The sound volume ratio (L) corresponds to the sound volume ratio of the reflected sound arrival time sound volume to the direct sound arrival time sound volume.
Naturally, the threshold values illustrated in FIG. 15 are examples, and the threshold values are not limited to the examples in FIG. 15. Furthermore, in FIG. 15, mainly threshold values when the angle (θ) of the direct sound arrival direction is 0 degrees are exemplified. However, threshold values when the direct sound arrival direction (θ) is not 0 degrees are also stored in memory 1404.
Further, in the above description, the threshold values are stored in an arrangement that has the angle (θ) of the direct sound (more specifically, the angle (θ) of the direct sound arrival direction) and the angle (γ) of the reflected sound (more specifically, the angle (γ) of the reflected sound arrival direction) as independent variables or indexes. However, the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction need not be used as independent variables.
For example, the angular difference between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be used. This angular difference corresponds to the angle formed between the direct sound arrival direction and the reflected sound arrival direction, and may be expressed as the arrival angle between a direct sound and a reflected sound.
FIG. 16 is a diagram illustrating relationships between angular differences, time differences, and threshold values. For example, threshold values pre-calculated by using, as a variable, the angular difference (φ) between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be stored as in the example illustrated in FIG. 16. Naturally, the threshold values illustrated in FIG. 16 are examples, and the threshold values are not limited to the examples in FIG. 16.
In the example in FIG. 16, the number of variables used for deriving threshold values may be reduced. Thus, it is possible to reduce the number of threshold values stored in memory 1404. Therefore, it is possible to decrease the amount of data stored in memory 1404.
Furthermore, when the angular difference (φ) between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction is used, the threshold value data may be stored in a two-dimensional arrangement. Moreover, in the selection processing, the difference between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be calculated by using a three-dimensional arrangement.
The method for selecting reflected sounds using threshold values based on the directions of arrival will be described later.
[First Variation of Threshold Value Setting Method]
In the examples in FIG. 12A, FIG. 12B, and FIG. 12C, threshold values in a plurality of formats and of a plurality of types may be stored in spatial information managers 1201 and 1211. Then, of the threshold values having a plurality of formats and a plurality of types, the format and the type of the threshold values to be used in the selection processing of the reflected sounds may be decided. Specifically, as illustrated in Example 3 of FIG. 12C, in the time differences (T) corresponding to the reflected sound arrival times, the largest threshold value may be adopted.
Moreover, as illustrated in Example 4, the masking threshold value, the echo detection limit threshold value, and a threshold value indicating the minimum sound volume for reproduction in the virtual space may be stored. Then, the largest threshold value at the time difference (T) corresponding to the reflected sound arrival time may be adopted.
[Second Variation of Threshold Value Setting Method]
As another example of the threshold value setting method, a method for setting threshold values in accordance with a characteristic of direct sounds will be described.
FIG. 17 is a block diagram illustrating another configuration example of renderer 1300 illustrated in FIG. 7. Renderer 1300 in FIG. 17 is different from renderer 1300 in FIG. 7 in the respect that renderer 1300 in FIG. 17 includes threshold value adjuster 1304. The description other than threshold value adjuster 1304 is the same as the matters described regarding FIG. 7, and has thus been omitted.
Threshold value adjuster 1304 selects, from the threshold value data, threshold values that are to be used by selector 1302, based on information indicating a characteristic of an audio signal. Alternatively, threshold value adjuster 1304 may adjust the threshold values included in the threshold value data, based on the information indicating a characteristic of the audio signal.
The information indicating a characteristic of the audio signal may be included in the input signal. Then, threshold value adjuster 1304 may obtain the information indicating a characteristic of the audio signal from the input signal. Alternatively, analyzer 1301 may derive a characteristic of the audio signal by analyzing the audio signal included in the input signal accepted by analyzer 1301, and output the information indicating the characteristic of the audio signal to threshold value adjuster 1304.
The information indicating a characteristic of the audio signal may be obtained before starting the rendering processing, or may be obtained each time during rendering.
Furthermore, threshold value adjuster 1304 need not be included in audio signal processing device 1001; another transmission device may have the role of threshold value adjuster 1304. In this case, analyzer 1301 or selector 1302 may obtain, from the other transmission device via communication I/F 1403, the information indicating a characteristic of the audio signal, the threshold value data corresponding to the characteristic, or information for adjusting the threshold value data in accordance with the characteristic.
FIG. 18 is a flowchart illustrating another example of selection processing. FIG. 19 is a flowchart illustrating yet another example of selection processing. In FIG. 18 and FIG. 19, the threshold value is set in accordance with a characteristic of the direct sound. Specifically, in FIG. 18, threshold value adjuster 1304 identifies a threshold value from the threshold value data, based on the time difference (T) and a characteristic of the audio signal. In FIG. 19, threshold value adjuster 1304 adjusts, based on a characteristic of the audio signal, the threshold value identified from the threshold value data based on the time difference (T).
Hereinafter, the operations of each example will be described. Note that description has been omitted for processes that are shared with the example in FIG. 14.
First, an example of the processing illustrated in FIG. 18 will be described. Here, the threshold value data is stored beforehand in memory 1404 for each nature of direct sound. Accordingly, a plurality of items of threshold value data corresponding to a plurality of natures are stored beforehand in memory 1404. Then, threshold value adjuster 1304 identifies, from the plurality of items of threshold value data, the threshold value data to be used in the selection processing of reflected sounds.
For example, threshold value adjuster 1304 obtains a characteristic of a direct sound based on the input signal (S211). Threshold value adjuster 1304 may obtain a characteristic of the direct sound that is associated with the input signal. Threshold value adjuster 1304 may then identify the threshold value corresponding to the time difference (T) and the characteristic of the direct sound (S212).
Furthermore, as illustrated in FIG. 19, threshold value adjuster 1304 may adjust the threshold value identified by selector 1302, based on the characteristic of the direct sound (S221).
In any of these cases, the information indicating a characteristic of the audio signal, the information for adjusting the threshold value in accordance with the characteristic of the audio signal, or both of these may be included in the input signal. Threshold value adjuster 1304 may adjust the threshold value using one or both of these.
Furthermore, the information indicating a characteristic of the audio signal, the information for adjusting the threshold value, or both of these may be transmitted by another input signal aside from the input signal that includes the audio signal. In this case, information for associating the other input signal aside from the input signal may be included in the input signal that includes the audio signal, or information for associating the other input signal with the input signal may be stored in memory 1404 together with the information on threshold values.
In the examples in FIG. 18 and FIG. 19, the threshold value used in selecting each reflected sound is set in accordance with a characteristic of the direct sound, that is, a characteristic of the audio signal. Threshold value data preset for each characteristic may be used, as in FIG. 18, or the threshold value may be adjusted in accordance with the characteristics of the audio signal, as in FIG. 19. Furthermore, threshold value data parameters may be adjusted in accordance with the characteristic of the audio signal.
Moreover, the operations performed by threshold value adjuster 1304 may be performed by analyzer 1301 or selector 1302. For example, analyzer 1301 may obtain a characteristic of the audio signal. Furthermore, selector 1302 may set threshold values in accordance with the characteristic of the audio signal.
Next, the relationship between the characteristic of the audio signal and the threshold value will be described.
Two short sounds that arrive at the listener's ears in succession are heard as one sound if the time period interval between the two short sounds is sufficiently short. This phenomenon is referred to as the precedence effect. The precedence effect is known to only occur with respect to unconnected sounds, that is, transient sounds (NPL 1). Thus, when an audio signal indicates a stationary sound, the echo detection limit may be set lower than when the audio signal indicates a non-stationary sound.
In other words, the threshold value is set low in accordance with the characteristics of this precedence effect when, for example, a direct sound is a stationary sound. Furthermore, the threshold value may be set lower as the stationarity is greater.
An example of processing when the characteristic of the audio signal is stationary will be explained. First, threshold value adjuster 1304 or analyzer 1301 makes a determination on the stationarity based on the amount of variation in a frequency component of an audio signal accompanying the passage of time. For example, when the amount of variation is small, it is determined that the stationarity is high. Conversely, when the amount of variation is great, it is determined that the stationarity is low. As a result of the determination, a graph indicating the level of stationarity may be set, or a parameter indicating the stationarity in accordance with the amount of variation may be set.
Next, threshold value adjuster 1304 adjusts the threshold value data or the threshold values based on information indicating the stationarity, such as the graph or the parameter indicating the stationarity of the audio signal, and sets the adjusted threshold value data or threshold values as threshold value data or threshold values to be used by selector 1302.
Alternatively, a parameter for setting the threshold value data in accordance with the information indicating direct sound stationarity may be stored beforehand in memory 1404. In this case, threshold value adjuster 1304 may make a determination on the stationarity of the audio signal and set the threshold value data to be used in the selection of reflected sounds, based on the information indicating stationarity and the parameter.
Alternatively, a plurality of parameters for threshold value data may be stored beforehand in memory 1404, corresponding to a plurality of patterns of direct sound stationarity. In this case, threshold value adjuster 1304 may make a determination on the stationarity of the audio signal, select the threshold value data parameter based on the pattern of direct sound stationarity, and set the threshold value data to be used in the selection of reflected sounds, based on the threshold value data parameter.
Note that a determination on the stationarity of an audio signal may be made based on the amount of variation of the frequency component of the audio signal, each time an audio signal is inputted.
Alternatively, a determination on the stationarity of an audio signal may be made based on information indicating stationarity that is pre-associated with the audio signal. In other words, the information indicating audio signal stationarity may be associated with the audio signal and pre-stored in memory 1404. Analyzer 1301 may, each time an audio signal is inputted, obtain information indicating stationarity that is associated with the audio signal. Threshold value adjuster 1304 may then adjust the threshold values based on the information indicating stationarity that is associated with the audio signal.
As another example of threshold values being set in accordance with a characteristic of the audio signal, when an audio signal indicates short sounds (clicking sounds, etc.), the application scope of the echo detection limit may be set shorter than when an audio signal indicates long sounds. This processing is based on the characteristics of the precedence effect.
It is known that due to the precedence effect, two short sounds that arrive at the listener's ears in succession are heard as one sound if the time period interval between the two short sounds is sufficiently short. The upper limit of this time period interval is dependent on the length of the sounds. For example, the upper limit of this time period interval is about 5 ms for clicking sounds, but for complex sounds such as a human voice or music, the upper limit may be 40 ms (NPL 1).
In accordance with the characteristics of this precedence effect, for example, in the case of a sound for which the duration of a direct sound is short, threshold values for short time period lengths are set. Furthermore, threshold values for shorter time period lengths are set as the duration of the direct sound is shorter.
Threshold values for short time period lengths being set means that within a range in which the time difference (T) between a direct sound and a reflected sound is small, threshold values corresponding to an echo detection limit based on the characteristics of the precedence effect are set. Threshold values corresponding to the echo detection limit based on the characteristics of the precedence effect are not set outside of this range. In other words, outside of this range, threshold values are low. Thus, threshold values for short time period lengths being set for short sounds can correspond to low threshold values being set for short sounds.
As another example of threshold values being set in accordance with a characteristic of a direct sound, when a direct sound is an intermittent sound (such as speech), threshold values may be set lower than when a direct sound is a continuous sound (such as music).
For example, when a direct sound corresponds to speech, sound portions and silent portions repeat, and in the silent portions, only the post-masking effect occurs as the masking effect. On the other hand, when the direct sound is a continuous sound such as musical content, the masking effects that occur include both the post-masking effect and a simultaneous masking effect that results from sound occurring at that time. Consequently, the overall masking effect is greater in the case of music, etc. than in the case of speech, etc.
In accordance with masking effect characteristics such as those described above, threshold values may be set higher in the case of music, etc. than in the case of speech, etc. Conversely, threshold values may be set lower in the case of speech, etc. than in the case of music, etc. That is, threshold values may be set to be low when a direct sound has numerous intermittent portions.
As described above, the information indicating a characteristic of a direct sound may be information indicating the stationarity, intermittency, duration, etc. of the direct sound. Furthermore, the information indicating a characteristic of a direct sound may be any combination of these. Furthermore, the information indicating a characteristic of a direct sound may be information indicating the time variation of one of these, or may be information indicating the time variation of any combination of these. That is, the information indicating a characteristic of a direct sound may be information indicating the time variation of the direct sound.
For example, as indicated in the description of the stationarity determination, the information indicating a characteristic of a direct sound may be chronological data on the frequency characteristics. Here, the frequency characteristics may be expressed in a commonly used form such as a gain value per frequency band, a Fourier series with respect to a time axis signal, a linear predictive coding (LPC) coefficient or cepstral coefficient for determining a frequency envelope, or the like.
Furthermore, the information indicating a characteristic of a direct sound may be, as information indicating the intermittency of direct sound, information (an amplitude envelope outline) enumerating, in chronological order, a plurality of sets of: the duration for which the amplitude of a signal is stationary; and the amplitude value of the signal for that duration. Here, the amplitude value may be expressed as a ratio with respect to the reference sound volume.
Furthermore, the information indicating a characteristic of a direct sound may be information on the frequency characteristics of the direct sound. For example, the information indicating a characteristic of a direct sound may be information indicating the stationarity of the frequency characteristics of the direct sound. Specifically the information indicating a characteristic of a direct sound may be information (a spectrogram outline) enumerating, in chronological order, a plurality of sets of: a duration for which the variation in frequency characteristics is small; and the frequency characteristics of the signal for the duration. Here, the sound volume used as the reference for the frequency characteristics may be the reference sound volume.
For example, the information indicating the time variation of a direct sound may be information indicating a direct sound envelope. The information indicating the time variation of a direct sound may be used when the “minimum audible limit” described in [Example 4] of FIG. 12C is the threshold value. The signal compared to the minimum audible limit is the sound volume of a reflected sound.
The sound volume of a reflected sound is obtained by geometrical calculation from information on the positions of the sound source, the listener, and the reflection object. Specifically, the reference sound volume of the reflected sound is obtained with respect to the reference sound volume of the sound source. By adjusting the reference sound volume of the reflected sound using, as the information on a characteristic of a direct sound, information on the sound magnitude transition of the sound source, the sound volume of the reflected sound at each moment can be accurately determined. This is because the variation in the sound volume of the sound source is reflected in the variation in the sound volume of the reflected sound.
By comparing the sound volume of a reflected sound with the threshold value after adjusting the sound volume of the reflected sound, the reflected sounds that are auditorily necessary can be appropriately selected with more accuracy.
It goes without saying that naturally, the same result can be obtained by, without adjusting the reference sound volume of a reflected sound, adjusting the threshold value based on the inverse of the information on the sound magnitude transition of the sound source, and comparing the adjusted threshold value to the reference sound volume of the reflected sound. That is, the reference sound volume of a reflected sound may be adjusted using the information on the sound magnitude transition of the sound source, or the threshold value may be adjusted using the information on the sound magnitude transition of the sound source. The adjustment of the reference sound volume of a reflected sound and the adjustment of a threshold value correspond to each other.
Depending on the composition of the surface of an object that reflects sound, the reflectance of sound (the attenuation rate accompanying the reflection) is different for each frequency band. Accordingly, as described later, the reflectance (attenuation rate) of sound may be associated with each frequency band, for an object that reflects sound. Whether to select the reflected sound can be more accurately determined based on such reflectance information and spectrogram information. For example, processing such as the following is performed.
Specifically, for example, it is indicated by spectrogram information that a high-frequency component is more dominant than a low-frequency component at a certain time interval. Furthermore, for example, it is indicated by sound reflectance information that the reflectance is very low at the high-frequency component, compared to the low-frequency component.
In this case, there is a possibility that even if the amplitude of the sound source signal on the time axis is high, the sound volume of the reflected sound will be low, the sound volume of the reflected sound being obtained by multiplying, by the attenuation rate for each frequency band indicated by the reflectance information, the frequency component indicated by the spectrogram information. Thus, the reflected sound will not be selected.
As described above, the information indicating a characteristic of a direct sound may be information indicating the time variation of the direct sound. For example, the information indicating a characteristic of a direct sound may indicate a value obtained by analyzing the direct sound at a predetermined time length.
Specifically, the information indicating a characteristic of a direct sound may be information obtained by calculating the average energy or the average amplitude of the direct sound for each predetermined time length. Furthermore, the information indicating a characteristic of a direct sound may be information obtained by calculating the energy or average amplitude of the direct sound for each short time analysis length, and calculating the weighted average of the energy or the average amplitude for each long time analysis length that is longer than the short time analysis length.
More specifically, for example, the information indicating the time variation of a direct sound may be information obtained by calculating the energy or the average amplitude of the direct sound for each predetermined short time length (for example, 5 ms; hereinafter, time length frames are expressed as analysis frames). Furthermore, the information indicating the time variation of a direct sound may be information represented by a weighted average of the energy or average amplitude calculated using the past N−1 analysis frames.
When the energy of the n-th analysis frame is expressed by E(n), information I(n) indicating a characteristic of the direct sound is determined in accordance with the following formula.
Here, parameter a(i) denotes a weighting factor. Typically, a(i) is set such that a(i)≥0 and the sum of a(i) is 1. However, the method for setting a(i) is not limited thereto.
Note that information I(n) indicating a characteristic of direct sound is calculated each time 5 ms of direct sound is imported. That is to say, the time variation of information I(n) indicating a characteristic of a direct sound can be calculated with low delay. Thus, this method is suitably applied to an application requiring real-time performance.
Furthermore, information I(n) indicating a characteristic of a direct sound may be determined in accordance with the following formula.
Here, parameter b(i) denotes a weighting factor. Typically, b(i) is set such that b(i)≥0 and the sum of b(i) is 1. However, the method for setting b(i) is not limited thereto.
In this formula, information I(n) indicating a characteristic of a direct sound is recursively determined. Thus, the average energy for a long time length can be calculated with a small amount of computation.
Formula 1 and Formula 2 above can be considered to be filters in which E(n) is the input signal and I(n) is the output signal. In this case, Formula 1 is a moving average (MA) model filter and Formula 2 is an autoregressive (AR) model filter, and both of these have the characteristics of a low-pass filter. Furthermore, an ARMA model filter, in which both of these are combined, may be used.
Note that the method for deriving the information indicating the time variation of a direct sound is not limited to the above-described formulae or filters, and other well-known methods may be used. As described above, the information indicating the time variation of a direct sound indicates a value obtained by analyzing a direct sound at a predetermined time length. A direct sound may be analyzed from perspectives other than average energy.
Furthermore, as described above, the information indicating a characteristic of a direct sound may be information on the frequency characteristics of the direct sound. The information on the frequency characteristics of a direct sound may be information calculated by using the frequency characteristics of the direct sound. For example, the information on the frequency characteristics of a direct sound may be information obtained as the average energy of a low-frequency component, by averaging the low-frequency component of the direct sound at a predetermined analysis length.
Specifically, the low-frequency component of a direct sound can be determined by applying, to a direct sound contained in an analysis frame length, a filter having low-pass characteristics. From the energy or average amplitude of this low-frequency component, the information indicating a characteristic of a direct sound can be derived, similarly to Formula 1 described above.
When the energy of the low-frequency component of the n-th analysis frame is expressed by EL(n), information I(n) indicating a characteristic of the direct sound can be determined in accordance with the following formula.
Here, parameter c(i) denotes a weighting factor. Typically, c(i) is set such that c(i)≥0 and the sum of c(i) is 1. However, the method for setting c(i) is not limited thereto.
Note that information I(n) indicating a characteristic of direct sound is calculated each time 5 ms of direct sound is imported. That is to say, the time variation of information I(n) indicating a characteristic of a direct sound can be calculated with low delay. Thus, this method is suitably applied to an application requiring real-time performance.
Furthermore, similarly to Formula 2, information I(n) indicating a characteristic of a direct sound can be determined in accordance with the following formula.
Here, parameter d(i) denotes a weighting factor. Typically, d(i) is set such that d(i)≥0 and the sum of d(i) is 1. However, the method for setting d(i) is not limited thereto.
In this formula, information I(n) indicating a characteristic of a direct sound is recursively determined. Thus, the average energy for a long time length can be calculated with a small amount of computation.
Formula 3 and Formula 4 above can be considered to be filters in which E(n) is the input signal and I(n) is the output signal. In this case, Formula 3 is a moving average (MA) model filter and Formula 4 is an autoregressive (AR) model filter, and both of these have the characteristics of a low-pass filter. Furthermore, an ARMA model filter, in which both of these are combined, may be used.
In the above description, a filter having low-pass characteristics was used in the method for determining the low-frequency component of a direct sound, but the method for determining the low-frequency component of a direct sound is not limited thereto. Furthermore, the method for deriving the information indicating the time variation of direct sound is not limited to the above-described formulae or filters, and other well-known methods may be used. Furthermore, the spectrum of a direct sound can be calculated by applying frequency conversion to the direct sound. The energy or average amplitude of the low-frequency component of the spectrum can then be calculated.
Furthermore, in the above description, the MA model or the AR model is used for deriving the information indicating the time variation of a direct sound. The coefficients of these models may be predetermined fixed values, or may be variable values, which are values that temporally vary.
Furthermore, the relationship between the analysis frame length and the interval at which the information update threads are created may be as described below.
For example, when the time length of an analysis frame is TA (msec) and the interval at which information update threads are created is TU (msec), the value of N in (Formula 1) and (Formula 3), described above, in the MA filter may be approximately a value given by TU/TA. Furthermore, b(i) and d(i) (1≤i<N) in (Formula 2) and (Formula 4), described above, in the AR filter may be a value that results in the time constant of the filter being about TU (msec). The reason for this setting is that within the information update interval period, convergence of the filter is expected.
On the other hand, when, in the information indicating the time variation of a direct sound with the above-described setting, the value varies too sharply, I(n) may be precalculated. Furthermore, precalculated I(n) may be applied to the selection processing of reflected sounds. For example, in the processing of the t-th time frame, I(t+tau) may be used. Here, tau is a value defined in accordance with the characteristics of filter convergence. When the convergence is slow, the value of tau is large in comparison to when the convergence is fast.
Furthermore, as the information indicating the characteristics of a direct sound, information on auditory masking (frequency masking) calculated from the direct sound may be used. The auditory masking information indicates a threshold value of an amplitude value at a frequency region masked by a direct sound. Processing in which the amplitude values of reflected sounds in the same frequency range are compared to the threshold value, and a reflected sound having an amplitude value lower than the threshold value is not selected may be performed. The amplitude value of a reflected sound in a frequency range may be obtained by analyzer 1301 as information indicating the characteristics of a reflected sound.
When threshold values to be used in selecting reflected sounds are thus set in accordance with the characteristics of direct sound, it is possible to appropriately select reflected sounds that are auditorily necessary, and the characteristics of auditory sensitivity can be effectively reflected in three-dimensional sound reproduction system 1000. Processing to detect characteristics of direct sound, processing to determine threshold values in accordance with the characteristics, and processing to adjust the threshold values in accordance with the characteristics may be performed during the rendering processing, or may be performed before starting the rendering processing.
For example, these processes may be performed, for example, during virtual space creation (during software creation), when starting processing of the virtual space (when launching the software or starting rendering), or when there is an occurrence of an information update thread that periodically occurs in processing of the virtual space. Furthermore, the time of virtual space creation may be when the virtual space is built before starting acoustic processing, may be when information (spatial information) on the virtual space is obtained, or may be when software is obtained.
Here, in the information update thread, processing to update the spatial information managed by spatial information managers 1201 and 1211 is performed.
The role of the information update thread is, for example, processing to update, based on the position and orientation of the VR goggles worn by the listener, the position and orientation of the listener's avatar positioned in the virtual space, updating of the positions of objects that have moved within the virtual space, and the like. Such processing is covered within a processing thread that launches relatively infrequently, at approximately tens of Hz.
Processing to update the information indicating a characteristic of a direct sound may be performed by such a processing thread that is created infrequently. The reason for this is that characteristics of direct sound vary less frequently than audio processing frames for audio output occur. This makes it possible to relatively reduce the computational load of the processing. Furthermore, when the information is updated at an unduly high frequency, there is a risk of pulsive noise being generated. Updating the information at a low frequency also makes it possible to avoid such a risk.
[Third Variation of Threshold Value Setting Method]
As another example of a threshold value setting method, threshold values may be set in accordance with computation resources (CPU capability, memory resources, PC performance, remaining level of battery, etc.) for processing reproduction of the virtual space. More specifically, sensor 1405 of audio signal processing device 1001 detects the amount of computation resources, and when the amount of computation resources is low, the threshold values are set to be high. Since consequently, the sound volume of a greater number of reflected sounds falls below the threshold values, the number of reflected sounds on which binaural processing is to be performed can be reduced, whereby the amount of computation can be reduced.
Alternatively, when the signal processing is performed by equipment that is driven by a storage battery, such as a smartphone or VR goggles, it is expected that priority is given to allowing processing to be performed for a longer duration, and computation resources are used economically. In such a case, it is not necessary to detect the amount or remaining level of computation resources, and the threshold values may be set to be high.
[Fourth Variation of Threshold Value Setting Method]
As another example of a threshold value setting method, by including a threshold value setter, not illustrated, in audio signal processing device 1001 or audio presentation device 1002, threshold values can be set by the manager of the virtual space or the listener.
For example, an “energy-saving mode”, in which there are few reflected sounds to be heard and the amount of computation is low, or a “high-performance mode”, in which there are many reflected sounds to be heard and the amount of computation is high, may be selectable by the listener to whom audio presentation device 1002 is equipped. Alternatively, the mode may be selectable by the manager who manages three-dimensional sound reproduction system 1000 or by the creator of the three-dimensional sound content. Furthermore, not the mode, but the threshold values or the threshold value data may be directly selectable.
[First Variation of Operations of Renderer]
FIG. 20 is a flowchart illustrating a first variation of operations of audio signal processing device 1001. FIG. 20 illustrates the processing performed mainly by renderer 1300 of audio signal processing device 1001. In this variation, sound volume compensation processing is added to the operations of renderer 1300.
For example, analyzer 1301 obtains data (the input signal) (S301). Next, analyzer 1301 analyzes the data (S302). Next, selector 1302 determines whether to select reflected sounds based on the analysis results (S303). Next, reproducer 1303 performs sound volume compensation processing based on the reflected sounds that were not selected (S304). Next, reproducer 1303 performs acoustic processing on the direct sounds and the reflected sounds (S305). Reproducer 1303 then outputs the direct sounds and the reflected sounds as audio (S306).
The above-described processes (S301 to S306) other than the sound volume compensation processing (S304) are processes that are shared with the other examples described above; thus, explanation thereof has been omitted.
The sound volume compensation processing is performed in accordance with the reflected sounds that were not selected in the selection processing. For example, due to not selecting a reflected sound in the selection processing, an absence emerges in the sound volume sensation. The sound volume compensation processing reduces the incongruity that accompanies this absence in the sound volume sensation. As an example of compensating the sound volume sensation, the following two methods are disclosed. Either of these two methods may be used.
First, a method in which the sound volume sensation is compensated for by raising the sound volume of a direct sound will be described. Reproducer 1303 raises the sound volume of the direct sound by the amount of the sound volume of a reflected sound that was not selected, and generates the direct sound. Accordingly, the sound volume sensation lost due to the reflected sound not being generated is compensated for.
At the time of raising the sound volume, reproducer 1303 may raise the sound volume of each frequency component in accordance with the frequency characteristics of the reflected sound. In order to make such processing possible, an attenuation rate of the sound volume attenuated by the reflection object may be assigned to each of predetermined frequency bands. This makes it possible to derive the frequency characteristics of the reflected sound.
Next, a method in which the sound volume sensation is compensated for by causing a reflected sound to be synthesized in a direct sound will be described. In this method, reproducer 1303 adds, to a direct sound, a reflected sound that was not selected and generates the direct sound to compensate for the sound volume sensation that results from the reflected sound not being generated. The sound volume (amplitude), frequency, delay, and the like of the reflected sound that was not selected are reflected in the generated direct sound.
In the case of the method for raising the sound volume of the direct sound, while the amount of computation for the compensation processing is extremely slight, only the sound volume is compensated for. In the case of the method of causing a reflected sound to be synthesized in a direct sound, the amount of computation for the compensation processing is large compared to the method of raising the sound volume of the direct sound, but the characteristics of the reflected sound are more accurately compensated for.
Since in both cases, only the direct sound is generated and the reflected sound is not generated, the total amount of computation is reduced. In particular, since the amount of computation required for binaural processing, which includes processing to implement a head-related transfer function (HRTF), is reduced, the total amount of computation is greatly reduced. The reason for this is that the amount of computation required for binaural processing is much greater than the amount of processing required for the above-described compensation processing.
Note that when the reason for a reflected sound not being selected is that the sound volume of the reflected sound is less than the masking threshold value, the sound volume sensation is not lost; thus, the reflected sound may be simply removed without performing compensation processing.
[Second Variation of Operations of Renderer]
FIG. 21 is a flowchart illustrating a second variation of operations of audio signal processing device 1001. FIG. 21 illustrates the processing performed mainly by renderer 1300 of audio signal processing device 1001. In this variation, left-right sound volume difference adjustment processing is added to the operations of renderer 1300.
For example, analyzer 1301 analyzes the input signal (S401). Next, analyzer 1301 detects the direction of arrival of sounds (S402). Next, selector 1302 adjusts the difference in sound volume between the sounds perceived by the left and right ears (S403). Furthermore, selector 1302 adjusts the difference in the arrival time periods (delay) between the sounds perceived by the left and right ears (S404). Selector 1302 determines whether to select reflected sounds based on information on the adjusted sounds (S405).
The above-described processes (S401 to S405) other than the left-right sound volume difference adjustment processing (S403) and the delay adjustment processing (S404) are processes that are shared with the other examples described above; thus, explanation thereof has been omitted.
FIG. 22 is a diagram illustrating an arrangement example of an avatar, a sound source object, and an obstacle object. For example, in a case in which the front direction of the listener is 0 degrees, when, as in FIG. 22, the polarities (for example, positive-negative) of the direction of arrival (θ) of the direct sound and the direction of arrival (γ) of the reflected sound (the direction (γ) of the reflected sound) are different, the sound volume difference that occurs between the ears is corrected.
Specifically, when the polarities of θ and γ are different, the ear which mainly (first) perceives the sound is different for each of the direct sound and the reflected sound. In this case, as the left-right sound volume difference adjustment processing (S403), selector 1302 adjusts the sound volume of the direct sound in accordance with the position of the ear that mainly perceives the reflected sound. For example, by multiplying the sound volume when the direct sound arrives at the listener by (1.0−0.3 sin(θ))(0≤θ≤180), selector 1302 causes attenuation of the sound volume when the direct sound arrives at the listener.
By calculating the sound volume ratio of the sound volume of the reflected sound to the sound volume of the direct sound, corrected as described above, and comparing the calculated sound volume ratio with threshold values, selector 1302 determines whether to select reflected sounds. Accordingly, the sound volume difference that occurs between the ears is corrected, the sound volume of direct sounds that affect reflected sounds is more accurately derived, and the determination of whether to select reflected sounds is more accurately performed.
Furthermore, in addition to the left-right sound volume difference adjustment processing (S403), selector 1302 may, as delay adjustment processing (S404), delay the direct sound arrival time period in accordance with the positions of the ears at which a reflected sound is perceived. Specifically, selector 1302 may delay the direct sound arrival time period by adding, to the direct sound arrival time period, (a (sin θ+θ)/c) ms (where a is the radius of the head and c is the speed of sound).
[Third Variation of Operations of Renderer]
A method for setting threshold values in accordance with directions of arrival will be described.
FIG. 23 is a flowchart illustrating yet another example of the selection processing. Description has been omitted for processes that are shared with the example in FIG. 14. In the example in FIG. 23, selector 1302 selects reflected sounds by using threshold values in accordance with directions of arrival.
Specifically, from the direct sound arrival path (pd), the reflected sound arrival path (pr), and avatar orientation information D1, each calculated by analyzer 1301, selector 1302 calculates the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (the direction (γ) of the reflected sound), each defined using the orientation of an avatar as reference. In other words, selector 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The orientation of the avatar corresponds to the orientation of the listener. Avatar orientation information D1 may be included in the input signal.
By using three indexes including the time difference (T), in addition to the direct sound arrival direction (θ) and the reflected sound arrival direction (γ), selector 1302 identifies, from a three-dimensional arrangement such as that illustrated in FIG. 15, the threshold values to be used in the selection processing (S232).
As an example, a method for setting threshold values to be used in selection processing when, as in FIG. 22, an avatar, a sound source object, and an obstacle object are arranged will be described.
Position information on the avatar, the sound source object, and the obstacle object, and avatar orientation information D1 are obtained from the input information. The direction (θ) of the direct sound and the direction (γ) of the sound image of the reflected sound when the orientation of the avatar is determined to be 0 degrees are calculated by using these items of position information and orientation information D1. In the case of FIG. 22, the direction (θ) of the direct sound is about 20 degrees, and the direction (γ) of the sound image of the reflected sound is about 265 degrees (−95 degrees).
Next, referencing the threshold value data stored in the three-dimensional arrangement illustrated in FIG. 15, a threshold value is identified from an arrangement domain that corresponds to the values of the two directions (θ) and (γ), and the value of the time difference (T) calculated by analyzer 1301. When there is no index that corresponds to the values of (θ), (γ), and (T) that were calculated, the threshold value corresponding to the index that is closest may be identified.
As another method, threshold values may be identified by performing processing such as interpolation or extrapolation, based on one or more threshold values that correspond to one or more indexes that are closest to the values of (θ), (γ), and (T) that were calculated. For example, a threshold value corresponding to (20 degrees, 265 degrees, T) may be identified based on the four threshold values corresponding to the four indexes of (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).
Selection processing based on the difference between the direct sound arrival direction angle (θ) and the reflected sound arrival direction angle (γ) will be described.
For example, as illustrated in FIG. 16, threshold value data having, as a two-dimensional index arrangement: the angular difference (φ) between the direct sound arrival direction (θ) and the reflected sound arrival direction (γ); and the time difference (T) may be pre-created and set. In this case, the angular difference (φ) and the time difference (T) are referenced in the selection processing. Alternatively, the angular difference (φ) between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be calculated in the selection processing, and the angular difference (φ) calculated may be used to identify the threshold value.
Alternatively, threshold value data having, as an index arrangement, a combination of the angular difference (φ), the direct sound arrival direction (θ), and the time difference (T), or a combination of the angular difference (φ), the reflected sound arrival direction (γ), and the time difference (T) may be set.
Alternatively, as illustrated in FIG. 15, threshold value data having, as a three-dimensional index arrangement, values of (θ), (γ), and (T) may be set.
[Fourth Variation of Operations of Renderer]
The processing performed by the above-described analyzer 1301, selector 1302, and reproducer 1303 may, for example, be performed as pipeline processing as described in PTL 3.
FIG. 24 is a block diagram illustrating a configuration example for renderer 1300 to perform pipeline processing.
Renderer 1300 in FIG. 24 includes reverberation processor 1311, early reflection processor 1312, distance attenuation processor 1313, selector 1314, generator 1315, and binaural processor 1316. These constituent elements may be configured as a plurality of the constituent elements of renderer 1300 illustrated in FIG. 7, or may be configured as at least a part of the plurality of constituent elements of audio signal processing device 1001 illustrated in FIG. 5.
Pipeline processing refers to dividing the processing for applying acoustic effects into a plurality of processes and executing each of the plurality of processes one by one in order. The plurality of processes include, for example, signal processing on the audio signal, generation of parameters used for signal processing, and the like.
Renderer 1300 may perform reverberation processing, early reflection processing, distance attenuation processing, binaural processing, and the like as pipeline processing. However, these types of processing are examples, and the pipeline processing may include processes other than these, or may not include some of these processes. For example, the pipeline processing may include diffraction processing and occlusion processing. Furthermore, for example, the reverberation processing may be omitted when unneeded.
Furthermore, each process may be expressed as a stage. Moreover, the audio signals of the reflected sounds and the like generated as the result of the processes may be expressed as rendering items. The plurality of stages and the order of these stages in the pipeline processing are not limited to the example illustrated in FIG. 24.
Here, the parameters (the arrival paths, the arrival time periods, and the sound volume ratios related to direct sounds and reflected sounds) used in the selection processing are calculated in one of the plurality of stages for generating the rendering items. In other words, the parameters used for selecting the reflected sounds are calculated as a part of the pipeline processing for generating the rendering items. Note that it is not necessary for all of the stages to be performed by renderer 1300. For example, a part of the stages may be omitted, or may be performed by an element other than renderer 1300.
The reverberation processing, the early reflection processing, the distance attenuation processing, the selection processing, the generation processing, and the binaural processing that may be included as stages in the pipeline processing will be described. In each stage, the metadata included in the input signal may be analyzed, and the parameters used for generating the reflected sounds may be calculated.
In the reverberation processing, reverberation processor 1311 generates an audio signal indicating reverberation sound or the parameters used in generating the audio signal. Reverberation sound is a sound that arrives at the listener as reverberation after the direct sound. As one example, the reverberation sound is a sound that arrives at the listener at a relatively late stage (for example, approximately 100 to 200 ms after the arrival of the direct sound) after the early reflected sound (to be described later) arrives at the listener, and after undergoing more reflections (for example, several tens of times) than the early reflected sound.
Reverberation processor 1311 refers to the audio signal and spatial information included in the input signal, and calculates reverberation sound by using, as a function for generating reverberation sound, a predetermined function prepared beforehand.
Reverberation processor 1311 may generate reverberation sound by applying a known reverberation generation method to the audio signal included in the input signal. One example of a known reverberation generation method is the Schroeder method, but the known reverberation generation method is not limited to the Schroeder method. Furthermore, reverberation processor 1311 uses the shape and acoustic characteristics of a sound reproduction space indicated by the spatial information when applying the known reverberation generation method. In this way, reverberation processor 1311 can calculate parameters for generating reverberation sound.
In the early reflection processing, early reflection processor 1312 calculates parameters for generating early reflection sounds based on the spatial information. The early reflected sound is reflected sound that arrives at the listener at a relatively early stage (for example, approximately several tens of ms after the arrival of the direct sound) after the direct sound from the sound source object arrives at the listener, and after undergoing one or more reflections.
Early reflection processor 1312 references, for example, the audio signal and metadata, and calculates the path, from reflection objects, of reflected sound that arrives at the listener after being reflected by the reflection objects. For example, in calculation of the path, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the positions of reflection objects such as structures, the reflectance of reflection objects, and the like may be used.
Early reflection processor 1312 may calculate the path of the direct sound. The information of said path may be used as a parameter for early reflection processor 1312 to generate the early reflected sound, and may be used as a parameter for selector 1314 to select reflected sounds.
In the distance attenuation processing, distance attenuation processor 1313 calculates the sound volume of the direct sound and the reflected sound that arrive at the listener, based on the lengths of the paths of the direct sound and the reflected sound. The sound volume of the direct sound and the reflected sound that arrive at the listener attenuate, with respect to the sound volume of the sound source, in proportion to the distance of the path to the listener (in inverse proportion to the distance). Thus, distance attenuation processor 1313 is able to calculate the sound volume of the direct sound by dividing the sound volume of the sound source by the length of the direct sound path, and is able to calculate the sound volume of the reflected sound by dividing the sound volume of the sound source by the length of the path of the reflected sound.
In the selection processing, selector 1314 selects the reflected sounds to be generated, based on the parameters calculated before the selection processing. One of the selection methods of the present disclosure may be used for selection of the reflected sounds to be generated.
The selection processing may be performed on all of the reflected sounds, or may be performed only on the reflected sounds having high evaluation values based on the evaluation processing, as described above. In other words, the reflected sounds having low evaluation values may be determined as not selected, without performing the selection processing. For example, reflected sounds for which the sound volume is extremely low may be considered to be reflected sounds having low evaluation values, and may be determined as not selected.
Furthermore, for example, the selection processing may be performed on all of the reflected sounds. Then, the evaluation values of the reflected sounds selected in the selection processing may be determined, and the reflected sounds having low determined evaluation values may be redetermined as not selected.
The selection processing and the evaluation processing may each be performed independently, or may be performed in combination with each other. When the selection processing and the evaluation processing are performed in combination with each other, either of the two processes may be performed first.
In the generation processing, generator 1315 generates direct sounds and reflected sounds. For example, generator 1315 generates direct sounds based on the direct sound arrival times and arrival time sound volume, from the audio signal included in the input signal. Furthermore, for each reflected sound selected in the selection processing, generator 1315 generates the reflected sound based on the reflected sound arrival time and the arrival time sound volume, from the audio signal included in the input signal.
In the binaural processing, binaural processor 1316 performs signal processing so that the audio signal of the direct sound is perceived as sound arriving at the listener from the direction of the sound source object. Furthermore, binaural processor 1316 performs signal processing so that the reflected sounds selected by selector 1314 are perceived as sounds arriving at the listener from the reflection object.
For example, based on the position and orientation of the listener in the sound space, binaural processor 1316 performs processing to apply an HRIR DB so that sound arrives at the listener from the position of the sound source object or the position of the obstacle object.
Note that Head-Related Impulse Response (HRIR) is the response characteristic when one impulse is generated. Specifically, HRIR is the response characteristic obtained by converting from an expression in the frequency domain to an expression in the time domain by Fourier transforming the head-related transfer function, in which the change in sound caused by surrounding objects including the auricle, the head, and the shoulders is expressed as a transfer function. The HRIR DB is a database including such information.
Furthermore, the position and orientation of the listener in the sound space are, for example, the position and orientation of a virtual listener in a virtual sound space. The position and orientation of the virtual listener in the virtual sound space may change in accordance with movement of the head of the listener. Furthermore, the position and orientation of the virtual listener in the virtual sound space may be determined based on information obtained from sensor 1405.
The program(s), spatial information, HRIR DB, threshold value data, other parameters, and/or the like used in the above-described processing are obtained from memory 1404 included in audio signal processing device 1001, or from outside of audio signal processing device 1001.
Furthermore, the pipeline processing may contain other processes. Moreover, renderer 1300 may contain a processor that is not illustrated, for performing another process included in the pipeline processing. For example, renderer 1300 may include a diffraction processor and an occlusion processor.
The diffraction processor executes processing to generate an audio signal indicating sound including diffracted sound caused by an obstacle object between the listener and the sound source object in a three-dimensional sound field (space). Diffracted sound is sound that when an obstacle object is present between the sound source object and the listener, arrives at the listener from the sound source object by going around the obstacle object.
The diffraction processor references, for example, the audio signal and metadata, and calculates the path by which diffracted sound arrives at the listener from the sound source object by detouring around the obstacle object, and generates diffracted sound based on the calculated path. In the calculation of the path, the sound source object in the three-dimensional sound field (space), the positions of the listener and the obstacle object, the shape and size of the obstacle object, and the like may be used.
When a sound source object is present on the other side of an obstacle object, the occlusion processor generates an audio signal for a sound that passes from the sound source object through the obstacle object and is audible therethrough, based on the spatial information and information on the material, etc. of the obstacle object.
[Example of Sound Source Object]
As described above, in the position information assigned to the sound source object, a “point” in the virtual space indicates the position of a sound source object. In other words, as described above, the sound source is defined as a “point sound source”.
On the other hand, a sound source in a virtual space may be defined as an object that has a length, size, shape, and the like, i.e., as a sound source that is not a point sound source, but a spatially extended sound source. In this case, the distance between the listener and the sound source, and the direction of arrival of the sound are not determined. Consequently, reflected sounds originating from such a sound source may be limited to being selected by selector 1302 without performing analysis by analyzer 1301, or regardless of the analysis result. By doing so, it is possible to avoid the sound quality degradation that might occur by not selecting the reflected sound.
Alternatively, a representative point such as the center of gravity of the object may be determined, and the processing of the present disclosure may be applied on the assumption that sound is generated from that representative point. In this case, the threshold value may be adjusted in accordance with information on the spatial extension of the sound source.
[Examples of Direct Sound and Reflected Sound]
For example, direct sound is sound that has not been reflected by a reflection object, and reflected sound is sound that has been reflected by a reflection object. Direct sound may be sound that has arrived at the listener from a sound source without being reflected by a reflection objection, and reflected sound may be sound that has arrived at the listener from a sound source due to being reflected by a reflection object.
Furthermore, each of direct sound and reflected sound are not limited to being sound that has arrived at the listener, and may each be sound that will arrive at the listener. For example, direct sound may be sound that has been outputted from a sound source, or to put it differently, a sound source sound.
FIG. 25 is a diagram illustrating transmission and diffraction of a sound. As illustrated in FIG. 25, a direct sound may not arrive at the listener due to the presence of an obstacle object between the sound source object and the listener. In this case, a sound that arrives at the listener after being emitted from the sound source object and passing through the obstacle object may be considered to be a direct sound. Furthermore, a sound that arrives at the listener after being emitted from the sound source object and diffracted by the obstacle object may be considered to be a reflected sound.
Furthermore, the two sounds compared in the selection processing are not limited to a direct sound and a reflected sound based on sound emitted from one sound source. For example, the selection of a sound may be performed by performing a comparison between two reflected sounds based on a sound emitted from one sound source. In this case, the direct sound in the present disclosure may be understood to be the sound that reaches the listener first, and the reflected sound in the present disclosure may be understood to be the sound that reaches the listener afterward.
[Example of Bitstream Structure]
The bitstream includes, for example, an audio signal and metadata. The audio signal is sound data in which sound is expressed, and indicates, e.g., information on the frequency and intensity of sound. Furthermore, metadata includes spatial information on the sound space, which is the space of the sound field.
For example, the spatial information is information on the space in which the listener who hears sound based on the audio signal is positioned. Specifically, the spatial information is information about a predetermined position (localization position) in the sound space (for example, a three-dimensional sound field) for localizing the sound image of the sound at that predetermined position, that is, for causing the listener to perceive the sound as arriving from a direction that corresponds to the predetermined position. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.
The sound source object information is information on a sound source object that generates sound based on the audio signal. In other words, the sound source object information is information on an object (a sound source object) that reproduces the audio signal, and is information on a virtual sound source object located in a virtual sound space. Here, the virtual sound space may correspond to real-world space in which an object that generates sound is located, and the sound source object in the virtual sound space may correspond to an object that generates sound in a real-world space.
The sound source object information may indicate, for example, the position of the sound source object located in the sound space, the orientation of the sound source object, the directivity of the sound emitted by the sound source object, whether the sound source object belongs to an animate thing, whether the sound source object is a mobile body, and the like. For example, the audio signal is associated with one or more sound source objects indicated by the sound source object information.
The bitstream includes, for example, metadata (control information) and an audio signal.
The audio signal and metadata may be contained in a single bitstream or may be separately contained in a plurality of bitstreams. Furthermore, the audio signal and metadata may be contained in a single file or may be separately contained in a plurality of files.
The bitstream may exist for each sound source or may exist for each playback time. Even in a case in which bitstreams exist for each playback time, a plurality of bitstreams may be processed in parallel simultaneously.
Metadata may be assigned to each bitstream, or may be collectively assigned to a plurality of bitstreams as information for controlling the plurality of bitstreams. In this case, the plurality of bitstreams may share the metadata. Furthermore, the metadata may be assigned for each playback time.
When a plurality of bitstreams or a plurality of files exist, information indicating a relevant bitstream or a relevant file may be contained in one or more bitstreams or one or more files.
Alternatively, information indicating a relevant bitstream or a relevant file may be contained in each of all of the bitstreams or each of all of the files.
Here, the relevant bitstream or the relevant file is, for example, a bitstream or file that may be used simultaneously during acoustic processing. Furthermore, a bitstream or file that collectively describes the information indicating the relevant bitstream or the relevant file may be included.
Here, the information indicating the relevant bitstream or the relevant file may be, for example, an identifier indicating a relevant bitstream or a relevant file. Furthermore, the information indicating the relevant bitstream or the relevant file may be, for example, a file name indicating a relevant bitstream or a relevant file, a uniform resource locator (URL), a uniform resource identifier (URI), or the like.
In this case, an obtainer identifies and obtains a relevant bitstream or a relevant file based on the information indicating the relevant bitstream or the relevant file. Furthermore, the information indicating the relevant bitstream or the relevant file may be included in a bitstream or a file, and the information indicating the relevant bitstream or the relevant file may be included in a different bitstream or a different file.
Here, the file including the information indicating the relevant bitstream or the relevant file may be, for example, a control file such as a manifest file used in content distribution.
Note that the entire metadata or part of the metadata may be obtained from somewhere other than a bitstream of the audio signal. For example, either one of metadata for controlling an acoustic sound or metadata for controlling a video may be obtained from somewhere other than from a bitstream, or both may be obtained from somewhere other than from a bitstream.
Furthermore, the metadata for controlling a video may be included in the bitstream obtained by three-dimensional sound reproduction system 1000. In this case, three-dimensional sound reproduction system 1000 may output the metadata for controlling a video to a display device that displays images or a stereoscopic video reproduction device that reproduces stereoscopic videos.
[Examples of Information Included in Metadata]
The metadata may be information used for describing a scene expressed in the sound space. As used herein, the term “scene” refers to a collection of all elements that represent three-dimensional video and acoustic events in the sound space, which are modeled in three-dimensional sound reproduction system 1000 using metadata.
Thus, the metadata may include not only information for controlling acoustic processing, but also information for controlling video processing. The metadata may include only one among the information for controlling acoustic processing or the information for controlling video processing, or may include both.
Three-dimensional sound reproduction system 1000 generates virtual acoustic effects by performing acoustic processing on the audio signal using the metadata included in the bitstream and additionally obtained interactive listener position information. Early reflection processing, obstacle processing, diffraction processing, occlusion processing, and reverberation processing may be performed as acoustic effects, and other acoustic processing may be performed using the metadata. For example, an acoustic effect such as a distance decay effect, localization, or a Doppler effect may be added.
In addition, information for switching between on and off of all or one or more of the acoustic effects, and priority information regarding a plurality of processes for the acoustic effects may be added to the metadata.
As an example, the metadata includes information about a sound space including a sound source object and an obstacle object and information about a localization position for localizing the sound image at a predetermined position in the sound space (that is, causing the listener to perceive the sound as arriving from a predetermined direction).
Here, an obstacle object is an object that can influence a sound emitted by a sound source object and perceived by the listener, by, for example, blocking or reflecting the sound between the sound source object and the listener. The obstacle object can include an animal or a movable body such as a machine, in addition to a stationary object. The animal may be a person or the like.
Furthermore, when a plurality of sound source objects are present in a sound space, another sound source object may be an obstacle object for a certain sound source object. In other words, non-sound-emitting objects such as building materials or inanimate objects, and sound source objects that emit sound can both be obstacle objects.
The metadata includes information indicating all or part of the shape of the sound space, the shapes and positions of obstacle objects in the sound space, the shapes and positions of sound source objects in the sound space, and the position and orientation of the listener in the sound space.
The sound space may be either a closed space or an open space. Furthermore, the metadata may include information indicating the reflectance of each obstacle object that can reflect sound in the sound space. For example, the floor, walls, ceiling, and the like constituting the boundaries of the sound space can be included in the obstacle objects.
The reflectance is an energy ratio between a reflected sound and an incident sound, and may be set for each sound frequency band. Of course, the reflectance may be uniformly set, irrespective of the sound frequency band. Note that when the sound space is an open space, for example, parameters such as a uniformly set attenuation rate, diffracted sound, and early reflected sound may be used.
The metadata may include information other than reflectance as a parameter with regard to an obstacle object or a sound source object. For example, the metadata may include information on the material of an object as a parameter related to both of a sound source object and a non-sound-emitting object. Specifically, the metadata may include information such as the diffusivity, transmittance, and sound absorption rate.
For example, information on a sound source object may include information indicating, for example, sound volume, radiation characteristics (directivity), a reproduction condition, the number and types of sound sources of one object, and a sound source region of an object. The reproduction condition may determine whether a sound is, for example, a sound that is continuously being emitted or is emitted at an event. The sound source region of an object may be determined by the relative relationship between the position of the listener and the position of the object, or may be determined using the object as a reference.
For example, when the sound source region is determined by the relative relationship between the position of the listener and the position of the object, it is possible to cause the listener to perceive sound E from the right side of the object and sound F from the left side of the object, the right side and the left side being as seen from the listener.
Furthermore, when the sound source region is determined using the object as a reference, it is possible to fix what sound is emitted from what region of the object, using the object as a reference. For example, it is possible, when the listener sees the object from the front, to cause the listener to perceive a high sound from the right side of the object and a low sound from the left side of the object. Furthermore, it is possible, when the listener sees the object from the rear, to cause the listener to perceive a low sound from the right side of the object and a high sound from the left side of the object.
Metadata related to the space may include the time period until early reflected sound, the reverberation time period, the ratio of direct sound to diffuse sound, and the like. When the ratio between a direct sound and a diffuse sound is zero, the listener can be caused to perceive only the direct sound.
BRIEF SUMMARY
Here, the present embodiment is briefly summarized.
When the relationship between a direct sound and a reflected sound is analyzed and the direct sound is the leading sound and the reflected sound is the lagging sound, if the relationship is such that the precedence effect occurs, i.e., if the reflected sound falls below the echo detection limit, the reflected sound is not perceived; thus, the auditory impact on the listener is small, even if the reflected sound is removed.
FIG. 26 is a diagram illustrating an example of the positional relationship between a listener and an obstacle object, according to the present embodiment. FIG. 27 is a diagram illustrating another example of the positional relationship between a listener and an obstacle object, according to the present embodiment. It should be noted that FIG. 26 has the same positional relationship as that illustrated in FIG. 9, and FIG. 27 has the same positional relationship as that shown in FIG. 10. Furthermore, FIG. 28 is an example of an echo detection limit threshold value according to the present embodiment. It should be noted that the echo detection limit threshold value shown in FIG. 28 is an example of the threshold value data shown in FIG. 12C and the like.
For example, comparing the positional relationship in FIG. 26 with the positional relationship in FIG. 27, the sound volume of the reflected sound heard by the listener at the positional relationship in FIG. 26 is lower than the sound volume of the reflected sound heard by the listener at the positional relationship in FIG. 27. This is because the path length until the reflected sound arrives in the positional relationship in FIG. 26 is longer than the path length until the reflected sound arrives in the positional relationship in FIG. 27.
Thus, when making the determination based solely on the sound volume of the reflected sound, the reflected sound illustrated in FIG. 26 has less auditory impact than the reflected sound illustrated in FIG. 27. However, comparing the arrival time of the reflected sound at the listening position illustrated in FIG. 26 with the arrival time of the reflected sound at the listening position illustrated in FIG. 27, the reflected sound arrives later in the case illustrated in FIG. 26.
Therefore, when the determination is made in terms of the echo detection limit, as illustrated in FIG. 28, the reflected sound illustrated in FIG. 27 is below the echo detection limit and is thus not perceived by the listener as a reflected sound, while the reflected sound illustrated in FIG. 26 is above the echo detection limit and is thus perceived by the listener as a reflected sound.
In the present embodiment, the amount of computation involved in processing reflected sounds is reduced by utilizing this to determine the auditory importance of reflected sounds, and not reproducing reflected sounds that are unimportant.
The above description is a brief summary of the present embodiment.
Here, attention is directed to sound volume ratio L and the threshold value.
The threshold value (echo detection limit threshold value) that is compared to sound volume ratio L is based on the auditory precedence effect. Therefore, as described above, when sound volume ratio L is calculated based solely on physical characteristics, the result of selecting whether to reproduce the reflected sound (the selection result) may not align with the listener's actual perception.
That is, without taking the auditory sensitivity of the listener into consideration, whether to output an audio signal indicating the reflected sound (more specifically, an output signal based on that audio signal) is selected, and there are cases in which such an output signal is output and the listener hears the sound (the reflected sound) represented by that output signal. In such cases, the listener hears a sound that differs from his/her own auditory perception, leading to a sense of incongruity.
Therefore, the following is a more detailed description of an audio signal processing method that can appropriately reduce the amount of computation and the computational load in a sound space, while taking auditory sensitivity into consideration.
Embodiment 2
Embodiment 2 is described below. The description below is centered on the points of difference from Embodiment 1, and descriptions of points in common are omitted or simplified.
[Configuration of Renderer]
First, the configuration of renderer 2300 according to the present embodiment is described. FIG. 29 is a block diagram illustrating a configuration example of renderer 2300 according to the present embodiment.
Renderer 2300 includes analyzer 2301, selector 2302, and reproducer 2303. It should be noted that as described above, the audio signal processing device according to the present embodiment is an example of a decoding device. The decoding device includes a decoder, and the decoder includes renderer 2300. In other words, it can be stated that the audio signal processing device according to the present embodiment includes analyzer 2301, selector 2302, and reproducer 2303. Renderer 2300 applies acoustic processing to sound data included in the input signal, and outputs the result.
Similarly to Embodiment 1, the input signal includes, for example, spatial information, sensor information, and sound data. The spatial information also includes physical information such as reflection coefficients, transmission coefficients, and diffraction coefficients of non-emitting objects (obstacle objects).
It should be noted that in the present embodiment, mainly reflected sound, which is an example of indirect sound, is used for description, but the same processing is performed even if indirect sound other than reflected sound is used. Furthermore, examples of indirect sound include reflected sound, diffracted sound, and the like.
Analyzer 2301 may be able to perform all of the processing or some of the processing performed by analyzer 1301 according to Embodiment 1.
Similarly to analyzer 1301 according to Embodiment 1, analyzer 2301 performs analysis on the audio signal included in the input signal, as well as the spatial information received from spatial information managers 1201 and 1211. Analyzer 2301 thus calculates the information necessary to generate direct sound and reflected sound with reproducer 2303, as well as the information necessary to select whether to generate reflected sound. The method by which analyzer 2301 calculates these items of information is as described in Embodiment 1.
Analyzer 2301 also performs analysis processing on the input signal, as performed by analyzer 1301 according to Embodiment 1, in S101 of FIG. 8. In other words, analyzer 2301 analyzes the input signal input to the audio signal processing device according to the present embodiment to detect direct sound and reflected sound that may be generated in the sound space.
When such direct sound and reflected sound are detected, analyzer 2301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data.
More specifically, analyzer 2301 creates an audio signal indicating reflected sound and an audio signal indicating direct sound based on: the position information of the sound source objects, the position information of the non-sound emitting objects (obstacle objects), and the position information and physical information of the listener that are included in the spatial information; and the sound data.
Specifically, analyzer 2301 creates an audio signal generated within the virtual space, based on the spatial information and the sound data, and then assigns attribute information indicating an attribute identifying the audio signal to the created audio signal, thereby creating an audio signal containing the attribute information. An audio signal including attribute information is created for each sound generated in the virtual space. The attribute includes information indicating whether the sound indicated by the audio signal is a direct sound or a reflected sound (indirect sound). As an example, in the present embodiment, the attribute is information indicating whether the sound indicated by the relevant audio signal is direct sound or reflected sound. Furthermore, the attribute information may include information necessary to radiate the audio signal into the sound space, such as, for example, gain information, gain characteristic information for each frequency bandwidth, position information, directivity information, and the like. In other words, the relevant necessary information may be retained in the attribute information. Furthermore, attribute information may be tied to the audio signal as metadata. The gain characteristics for each frequency bandwidth of the audio signal included in the attribute information may be identified based on, for example, the spatial information included in the input information. Information indicating frequency characteristics that indicate auditory sensitivity may be identified based on, for example, the spatial information included in the input information, especially as information tied to the avatar of the listener.
It should be noted that for simplicity, an audio signal for which the attribute is information indicating reflected sound (indirect sound) may be described as an audio signal indicating reflected sound (indirect sound), and an audio signal for which the attribute is information indicating direct sound may be described as an audio signal indicating direct sound.
Sound that directly reaches the listener's head from one sound source is direct sound, and sound that reaches the listener's head after being output from that one sound source and then reflected by a non-emitting object or diffracted by a non-emitting object is indirect sound (reflected sound or diffracted sound).
It should be noted that in the present embodiment, analyzer 2301 creates an audio signal indicating a reflected sound (indirect sound) and an audio signal indicating the direct sound associated with that reflected sound (indirect sound).
Furthermore, a direct sound associated with an indirect sound means a direct sound that originates from the same sound source as that indirect sound. An indirect sound associated with a direct sound means an indirect sound that originates from the same sound source as that direct sound. Furthermore, a reflected sound is a sound resulting from the direct sound associated with that reflected sound being reflected by a reflector.
An audio signal for which the attribute is information indicating a reflected sound (indirect sound) includes information indicating an audio signal for the direct sound associated with that reflected sound (indirect sound).
Analyzer 2301 may cause the created audio signal to be stored in memory included in analyzer 2301. Furthermore, analyzer 2301 also creates a plurality of audio signals and causes the plurality of audio signals to be stored in the memory.
Furthermore, as in Embodiment 1, analyzer 2301 may calculate, for each of the direct sound and reflected sound, values related to: the path until arriving at the listening position; the time period taken until arrival; the sound volume at arrival; and the like. Similarly, analyzer 2301 may calculate information indicating the relationship between the direct sound and the reflected sound, such as, for example, a value related to the time difference between when the direct sound arrives and when the reflected sound arrives (the time difference between the direct sound and the reflected sound), and the like.
It should be noted that the reflected sound arrival time sound volume (Ir) and the direct sound arrival time sound volume (Id) are examples of the arrival time sound volume and the like. The direct sound arrival time sound volume (Id) refers to the sound volume of a direct sound at the time of arrival at the listening position, which is the location at which the listener is present in a virtual space. In other words, the direct sound arrival time sound volume (Id) is the sound volume of the direct sound at the listening position. The reflected sound arrival time sound volume (Ir) refers to the sound volume of a reflected sound, which is an example of an indirect sound, at the time of arrival at the listening position. In other words, the reflected sound arrival time sound volume (Ir) is the sound volume of the indirect sound (the sound volume of the reflected sound) at the listening position.
In the present embodiment, each of the audio signal whose attribute is information indicating reflected sound and the audio signal whose attribute is information indicating direct sound may include information indicating the sound volume, at the listening position, of the sound indicated by the audio signal. In other words, in the present embodiment, the audio signal whose attribute is information indicating reflected sound includes information indicating the reflected sound arrival time sound volume (Ir) as the sound volume of the indirect sound (sound volume of the reflected sound). Similarly, the audio signal whose attribute is information indicating direct sound includes information indicating the direct sound arrival time sound volume (Id) as the sound volume of the direct sound.
It should be noted that the audio signal whose attribute indicates reflected sound may also include information indicating the sound volume of the indirect sound (sound volume of the reflected sound) and the sound volume of the direct sound associated with that reflected sound. Similarly, the audio signal whose attribute is information indicating direct sound may include information indicating the sound volume of the direct sound and the sound volume of the indirect sound (sound volume of the reflected sound) associated with that direct sound.
Selector 2302 may be able to perform all of the processing or some of the processing performed by selector 1302 according to Embodiment 1. Furthermore, selector 2302 determines whether the output signal based on the audio signal created by analyzer 2301 is to be output (reproduced) by reproducer 2303. That is, selector 2302 first specifies one audio signal from among the plurality of audio signals created by analyzer 2301 (for example, the audio signals indicating reflected sounds), and then selects whether reproducer 2303 is to generate and output an output signal based on the one audio signal specified.
Selector 2302 has obtainer 2302a, first calculator 2302b, second calculator 2302c, and selection processor 2302d.
Obtainer 2302a obtains an audio signal that includes attribute information and was created by analyzer 2301 and stored in the memory of analyzer 2301. Obtainer 2302a obtains, for example, an audio signal indicating a reflected sound (indirect sound) and an audio signal indicating a direct sound associated with the reflected sound (indirect sound). Furthermore, obtainer 2302a also obtains values related to the time difference between the direct sound and the reflected sound calculated by analyzer 2301.
First calculator 2302b calculates the first sound volume based on the audio signal that indicates the indirect sound and was obtained by obtainer 2302a. Here, first calculator 2302b calculates the first sound volume that is based on the sound volume of the indirect sound, which is the sound volume of the indirect sound indicated by the audio signal at the time at which the indirect sound arrives at the listening position that is the position at which the listener is present.
More specifically, first calculator 2302b calculates the first sound volume that is based on the sound volume of the reflected sound (that is, the reflected sound arrival time sound volume (Ir)), which is the sound volume of the reflected sound indicated by the audio signal at the time at which the reflected sound arrives at the listening position that is the position at which the listener is present.
Second calculator 2302c calculates the second sound volume based on the audio signal that indicates the direct sound and was obtained by obtainer 2302a. Here, second calculator 2302c calculates the second sound volume that is based on the sound volume of the direct sound (in other words, the direct sound arrival time sound volume (Id)), which is the sound volume of the direct sound indicated by the audio signal at the time at which the direct sound arrives at the listening position that is the position at which the listener is present.
It should be noted that in the present embodiment, the second sound volume is a sound volume different from the sound volume of the direct sound, but the sound volume of the direct sound may be used as the second sound volume, as-is.
Furthermore, selection processor 2302d selects whether reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound (indirect sound), based on the sound volume ratio between the second sound volume calculated and the first sound volume calculated.
When selection processor 2302d selects that reproducer 2303 is to output the output signal, selector 2302 outputs the audio signal obtained to reproducer 2303.
Reproducer 2303 may be able to perform all of the processing or some of the processing performed by reproducer 1303 according to Embodiment 1. Furthermore, reproducer 2303 obtains the audio signal output from selector 2302 and outputs an output signal that is based on the audio signal obtained.
Reproducer 2303 generates and outputs the output signal by performing binaural filtering processing and/or the like on the audio signal obtained. Binaural filtering processing is realized, for example, by subjecting the audio signal obtained to processing using a head-related transfer function.
Furthermore, reproducer 2303 may synthesize and output the audio signal that indicates the direct sound and was obtained by obtainer 2302a and the output signal generated.
Furthermore, reproducer 2303 may also generate and output an output signal by performing both binaural filter processing and diffusion filter processing on the audio signal output from selector 2302. The diffusion filter processing is, for example, processing that improves the realism of indirect sound by diffusing the reflected sound (indirect sound) indicated by the audio signal in the audio signal obtained. Furthermore, the diffusion filter processing is processing in which a filter is used to simulate the auditory intensity of the sound diffusion indicated by the audio signal obtained (i.e., to simulate the auditory intensity of the sound diffusion as perceived by the listener). Finite impulse filters and/or infinite impulse filters are used in the diffusion filter processing.
Hereinafter, an example of the operation of the audio signal processing method performed by the audio signal processing device according to the present embodiment (more specifically, renderer 2300) is described.
[Operation Example of Renderer]
FIG. 30 is a flowchart illustrating an operation example of the audio signal processing device according to the present embodiment. FIG. 30 illustrates the processing performed mainly by renderer 2300 included in the audio signal processing device according to the present embodiment. It should be noted that here, descriptions of common points with FIG. 8 according to Embodiment 1 are omitted or simplified.
First, analyzer 2301 performs analysis processing to analyze the input signal (S101a). More specifically, analyzer 2301 analyzes the input signal to detect direct sound and reflected sound that may be generated in the sound space. When such direct sound and reflected sound are detected, analyzer 2301 creates audio signals including attribute information, i.e., an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. Analyzer 2301 causes the created audio signal to be stored in the memory of analyzer 2301.
Analyzer 2301 analyzes the input signal and calculates, for each of the direct sound and reflected sound, values related to, e.g., the path until arrival at the listening position, the time taken to arrive, and the sound volume at the time of arrival, as well as a value related to the time difference between the direct sound and the reflected sound.
First, analyzer 2301 calculates the characteristics of each of the direct sound indicated by the audio signal created and the reflected sound indicated by the audio signal created. Specifically, the arrival time period and the arrival time sound volume when each of the direct sound and the reflected sound arrive at the listener (listening position) are calculated. It should be noted that the method shown in Embodiment 1 may be used as the method for calculating these arrival time periods and arrival time sound volumes.
It should be noted that as described above, the audio signal indicating reflected sound includes information indicating the reflected sound arrival time sound volume (Ir) as the sound volume of the reflected sound, and the audio signal indicating direct sound includes information indicating the direct sound arrival time sound volume (Id) as the sound volume of the direct sound.
Then, analyzer 2301 calculates the time difference (T) between the direct sound and the reflected sound (the time difference (T) between when the direct sound arrives and when the indirect sound arrives). It should be noted that the method shown in Embodiment 1 may be used as the method for calculating the time difference (T). Unlike step S101 according to Embodiment 1, the sound volume ratio (L) need not be calculated in step S101a.
Selector 2302 (more specifically, selection processor 2302d) performs the selection of reflected sounds (selection processing) (S102a). In other words, selector 2302 selects whether reproducer 2303 is to reproduce an output signal that is based on an audio signal that indicates a reflected sound and was created by analyzer 2301.
First, obtainer 2302a obtains an audio signal that includes attribute information and was created by analyzer 2301 and stored in the memory. Obtainer 2302a obtains, for example, at least one of an audio signal indicating a reflected sound (indirect sound) or an audio signal indicating a direct sound associated with the reflected sound (indirect sound). Here, obtainer 2302a obtains both. Furthermore, obtainer 2302a also obtains a value related to the time difference (T) between the direct sound and the reflected sound calculated by analyzer 2301.
As described above, the reflected sound and the direct sound associated with that reflected sound originate from the same sound source.
First calculator 2302b calculates the first sound volume that is based on the sound volume of the reflected sound, based on the audio signal that indicates the reflected sound and was obtained by obtainer 2302a.
Second calculator 2302c obtains the second sound volume that is based on the sound volume of the direct sound, based on the audio signal that indicates the direct sound and was obtained by obtainer 2302a.
Furthermore, selection processor 2302d selects whether reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound (indirect sound), based on: the sound volume ratio between the second sound volume calculated and the first sound volume calculated; and the time difference (T) between the direct sound and the reflected sound (indirect sound).
More specifically, selection processor 2302d selects that reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound (indirect sound), when the sound volume ratio is greater than or equal to the first threshold value determined according to the time difference (T) between the direct sound and the reflected sound (indirect sound).
The sound volume ratio is a value obtained by dividing the first sound volume, which is based on the reflected sound arrival time sound volume (Ir), by the second sound volume, which is based on the direct sound arrival time sound volume (Id).
Furthermore, the time difference (T) is the time difference (T) between: the direct sound associated with the reflected sound indicated by the audio signal obtained; and the reflected sound indicated by the audio signal obtained, and is the time difference (T) between the direct sound and the reflected sound, calculated by analyzer 2301 in step S101a. As described in Embodiment 1, the time difference (T) between a direct sound and a reflected sound is, for example, the time difference between the direct sound arrival time period (arrival time) and the reflected sound arrival time period (arrival time), but is not limited thereto.
The first threshold value is a value determined according to the time difference (T) between the direct sound associated with the reflected sound (indirect sound) and the reflected sound (indirect sound); in other words, the threshold value is a value dependent on the time difference (T) and is the value indicated by the threshold value data of Embodiment 1. The threshold value data is, for example, a graph having a horizontal axis that indicates the time difference (T) between a direct sound and reflected sounds and a vertical axis that indicates the sound volume ratios of reflected sounds to a direct sound, and is expressed as a threshold value (first threshold value) that demarcates whether each reflected sound is perceived.
More specifically, the threshold value data indicating the first threshold value is the data shown in FIG. 11 to FIG. 13 and the like.
Furthermore, in step S102a, obtainer 2302a obtains the gain characteristic of each predetermined frequency bandwidth related to the indirect sound (reflected sound) and the frequency characteristic indicating the auditory sensitivity. Similarly, in step S102a, obtainer 2302a obtains the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
The gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the frequency characteristic indicating the auditory sensitivity, and the gain characteristic of each predetermined frequency bandwidth related to the direct sound are stored, for example, in the memory of analyzer 2301. Obtainer 2302a obtains, from the memory, the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the frequency characteristic indicating the auditory sensitivity, and the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
It should be noted that the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the frequency characteristic indicating the auditory sensitivity, and the gain characteristic of each predetermined frequency bandwidth related to the direct sound may be obtained by obtainer 2302a via a communication line or the like.
Selector 2302 performs the selection processing as described above.
Next, the selection processing, particularly the calculation of the first sound volume and the second sound volume, is described in greater detail with reference to FIG. 31.
FIG. 31 is a flowchart illustrating an operation example of the selection processing according to the present embodiment. It should be noted that here, descriptions of common points with FIG. 14 according to Embodiment 1 are omitted or simplified.
First, selector 2302 specifies a reflected sound detected by analyzer 2301 (S201). In other words, obtainer 2302a of selector 2302 specifies the audio signal that includes the attribute information and was created by analyzer 2301 and stored in the memory, and obtains the audio signal specified. For example, obtainer 2302a specifies an audio signal indicating the reflected sound and obtains the audio signal specified. At this time, obtainer 2302a may also obtain an audio signal indicating the direct sound associated with that reflected sound (indirect sound).
Then, selector 2302 calculates the first sound volume that is based on the sound volume of the reflected sound (S210).
More specifically, first calculator 2302b of selector 2302 calculates the first sound volume that is based on the sound volume of the reflected sound, based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal indicating the reflected sound.
Furthermore, selector 2302 calculates the second sound volume that is based on the sound volume of the direct sound (S220).
More specifically, second calculator 2302c of selector 2302 calculates the second sound volume that is based on the sound volume of the direct sound, based on: a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound. Here, second calculator 2302c calculates the second sound volume that is based on the sound volume of the direct sound, based on the second correction characteristic and the audio signal indicating the direct sound. It should be noted that when the audio signal indicating the reflected sound includes information indicating the sound volume of the direct sound associated with the reflected sound, second calculator 2302c may calculate the second sound volume that is based on the sound volume of the direct sound, based on the second correction characteristic and the audio signal indicating the reflected sound.
Below, the calculations of the first sound volume and the second sound volume are described. The first sound volume and the second sound volume are calculated using Formula 5 and Formula 6 below.
First Sound Volume=GlobalGain_R*EqualiserGain_R (Formula 5)
Second sound volume=GlobalGain_D*EqualiserGain_D (Formula 6)
It should be noted that GlobalGain_R is the gain of the signal of the reflected sound across all frequency bands, and EqualiserGain_R is the representative value of the gain of the equalizer applied to the reflected sound for each frequency band. GlobalGain_D is the gain of the signal of the direct sound across all frequency bands, and EqualiserGain_D is the representative value of the gain of the equalizer applied to the direct sound for each frequency band.
The gain of the signal across all frequency bands may be calculated using the sound volume of the sound source of the sound and the distance from the sound source to the listener, as described in Embodiment 1. More specifically, the gain of the signal of the reflected sound across all frequency bands corresponds to the sound volume of the indirect sound (the sound volume of the reflected sound); in other words, it corresponds to the arrival time sound volume (Ir) included in the audio signal indicating the reflected sound. Furthermore, the gain of the signal of the direct sound across all frequency bands corresponds to the sound volume of the direct sound; in other words, it corresponds to the direct sound arrival time sound volume (Id) included in the audio signal indicating the direct sound.
It should be noted that when obtainer 2302a obtains only the audio signal indicating the reflected sound, the sound volume of the direct sound included in the audio signal indicating the reflected sound may be used as the gain of the signal of the direct sound across all frequency bands.
As described above, the representative value indicated by EqualiserGain_R is the representative value of the gain of the equalizer applied to the reflected sound for each frequency band. The gain of the equalizer applied to the reflected sound for each frequency band corresponds to the first correction characteristic described above. That is, EqualiserGain_R is the representative value of the first correction characteristic obtained by correcting, using the frequency characteristic indicating the auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the reflected sound.
For simplicity, the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound) may be referred to as the gain characteristic related to the reflected sound (indirect sound).
The gain characteristic related to the reflected sound (indirect sound) may be the gain, for each band, of an equalizer that realizes an adjustment in the frequency characteristic resulting from the fact that, for example, when a sound wave strikes an object (a non-sound-emitting object), the reflectance varies for each frequency component depending on the surface shape, material, or hardness of that object.
FIG. 32 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth related to a reflected sound according to the present embodiment. The gain characteristic related to reflected sound is indicated by the reflection coefficient, as shown in FIG. 32, but is not limited thereto.
FIG. 32 illustrates examples of gain characteristics corresponding to each of three types of wall surfaces. FIG. 32 shows the gain characteristic of each of a wall surface having a hard surface with numerous irregularities, a wall surface having a hard surface with few irregularities, and a wall surface having a soft surface. The gain characteristic corresponding to the object (non-sound-emitting object) generating the reflected sound is employed. It should be noted that the gain characteristics illustrated in FIG. 32 are merely examples.
Furthermore, while a ⅓-octave band is used here as the predetermined frequency bandwidth, this is not intended to be limiting.
The frequency characteristic indicating the auditory sensitivity is a frequency characteristic indicating the sound volume sensitivity of the listener, and for example, an A-weighting characteristic may be used. The A-weighting characteristic is a frequency-weighting characteristic that takes human hearing into consideration. FIG. 33A is a diagram illustrating a table showing a frequency characteristic indicating auditory sensitivity according to the present embodiment. FIG. 33B is a diagram illustrating the frequency characteristic sensitivity according to the present indicating the auditory embodiment. FIG. 33B is a diagram illustrating the frequency characteristic with the vertical axis expressed in decibels (dB). It should be noted that the frequency characteristic (A-weighting characteristic) indicating the auditory sensitivity illustrated in FIG. 33B displays a value for each ⅓-octave band, but this is not intended to be limiting.
Further, EqualiserGain_R is calculated as follows. Here, description is provided with reference to FIG. 34.
FIG. 34 is a diagram illustrating the gain characteristic related to the reflected sound according to the present embodiment, the frequency characteristic indicating the auditory sensitivity (A-weighting characteristic), and the first correction characteristic. FIG. 34 shows a value for each ⅓-octave band.
Using the above-described frequency characteristic (A-weighting characteristic) indicating the auditory sensitivity, the gain characteristic related to the reflected sound is corrected. Specifically, the value of the first correction characteristic corresponding to a given band is calculated by multiplying the value of the frequency characteristic indicating the auditory sensitivity for each ⅓-octave band, by the value of the gain characteristic related to the reflected sound corresponding to that band. In FIG. 34, since the vertical axis is represented on a logarithmic scale (dB), the value of the first correction characteristic is calculated by, for each frequency band, adding the value of the frequency characteristic indicating the auditory sensitivity and the value of the gain characteristic related to the reflected sound. Furthermore, when the vertical axis in FIG. 34 is expressed as a linear axis (scaling factor), it goes without saying that the calculation may be performed using multiplication.
The representative value of the first correction characteristic calculated in this manner is denoted by EqualiserGain_R. It should be noted that the representative value of the first correction characteristic may be, e.g., the average value, the maximum value, or the minimum value of the values of the first correction characteristic, or may be, e.g., the average value, the maximum value, or the minimum value of the first correction characteristic within a predetermined frequency band.
Furthermore, EqualiserGain_D is described.
As described above, the representative value indicated by EqualiserGain_D is the representative value of the gain of the equalizer applied to the direct sound for each frequency band. The gain of the equalizer applied to the direct sound for each frequency band corresponds to the second correction characteristic described above. That is, EqualiserGain_D is the representative value of the second correction characteristic obtained by correcting, using the frequency characteristic indicating the auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
For simplicity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound may be referred to as the gain characteristic related to the direct sound.
The gain characteristic related to the direct sound may, for example, be the gain of the equalizer for each band, indicating the gain of frequency components due to the state of the space through which the sound propagates. More specifically, the gain characteristic related to the direct sound may be the gain, for each band, of the equalizer realizing a frequency characteristic exhibiting a tendency such that higher frequency components experience greater gain attenuation due to factors such as the temperature or humidity of the space through which the sound propagates, or the density of fine particles like dust or pollen. FIG. 35 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth of a direct sound according to the present embodiment. The gain characteristic related to direct sound is not limited to that shown in FIG. 35.
As shown in FIG. 35, for the gain characteristic related to direct sound, a value is shown for each predetermined frequency bandwidth, that is, for each ⅓-octave band, but this is not intended to be limiting.
It should be noted that the frequency characteristic indicating the auditory sensitivity used to calculate EqualiserGain_D may also be the A-weighting characteristic illustrated in FIG. 33B.
Further, EqualiserGain_D is calculated as follows. Here, a description is provided with reference to FIG. 36.
FIG. 36 is a diagram illustrating the gain characteristic related to the direct sound, the frequency characteristic indicating the auditory sensitivity (A-weighting characteristic), and the second correction characteristic, according to the present embodiment. FIG. 36 illustrates the value for each ⅓-octave band.
Using the above-described frequency characteristic that indicates the auditory sensitivity (A-weighting characteristic), the gain characteristic related to the direct sound is corrected. Specifically, the value of the second correction characteristic corresponding to a given band is calculated by multiplying the value of the frequency characteristic indicating the auditory sensitivity for each ⅓-octave band, by the value of the gain characteristic related to the direct sound corresponding to that band. In FIG. 36, since the vertical axis is represented on a logarithmic scale (dB), the value of the second correction characteristic is calculated by, for each frequency band, adding the value of the frequency characteristic indicating the auditory sensitivity and the value of the gain characteristic related to the direct sound. Furthermore, when the vertical axis in FIG. 36 is expressed as a linear axis (scaling factor), it goes without saying that the calculation may be performed using multiplication.
The representative value of the second correction characteristic calculated in this manner is denoted by EqualiserGain_D. It should be noted that the representative value of the second correction characteristic may be, e.g., the average value, the maximum value, or the minimum value of the values of the second correction characteristic, or may be, e.g., the average value, the maximum value, or the minimum value of the second correction characteristic within a predetermined frequency band.
The first sound volume and the second sound volume are calculated using EqualiserGain_D and EqualiserGain_R calculated in this way, according to Formula 5 and Formula 6 described above.
That is, the first sound volume is calculated by multiplying the reflected sound arrival time sound volume (Ir) (the sound volume of the reflected sound) included in the audio signal indicating the reflected sound by the calculated EqualiserGain_R calculated. Similarly, the second sound volume is calculated by multiplying the direct sound arrival time sound volume (Id) (the sound volume of the direct sound) included in the audio signal indicating the direct sound by EqualiserGain_D calculated.
It should be noted that in the present embodiment, the A-weighting characteristic was used as the frequency characteristic indicating the auditory sensitivity, that is, the frequency characteristic indicating the sound volume sensitivity of the listener; however, this is not intended to be limiting. For example, the inverse characteristic of the equal loudness contour, the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position, or the like may be used as the frequency characteristic indicating the sound volume sensitivity of the listener.
FIG. 37 is a diagram illustrating the inverse characteristic of the equal loudness contour, which is another, first example of the frequency characteristic indicating the sound volume sensitivity of the listener according to the present embodiment. In FIG. 33B, which shows the A-weighting characteristic, the A-weighting characteristic exhibits a convex upward characteristic. Similarly, the inverse characteristic of the equal loudness contour exhibits a convex upward characteristic. Furthermore, the method of using the inverse characteristic of the equal loudness contour in the above-described correction may be the same as described above.
FIG. 38 is a diagram illustrating the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position, which is another, second example of the frequency characteristic indicating the sound volume sensitivity of the listener according to the present embodiment. Furthermore, the method of using the inverse frequency characteristic of the minimum audible angle of the sound source position in the above-described correction may be the same as described above.
Thus, in the present embodiment, the frequency characteristic indicating the sound volume sensitivity of the listener can be utilized as the frequency characteristic indicating the auditory sensitivity. Furthermore, as the frequency characteristic indicating the sound volume sensitivity, a frequency characteristic based on the inverse of the equal loudness contour, or the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position can be utilized.
Once more, a description is provided with reference to FIG. 31. Next, selection processor 2302d calculates the sound volume ratio between the second sound volume calculated and the first sound volume calculated (S202a).
Then, selection processor 2302d detects the time difference (T) between the direct sound and the reflected sound (S203). The time difference (T) has already been calculated by analyzer 2301 in step S101a. For example, data indicating the time difference (T) is stored in the memory of analyzer 2301, and selector 2302 detects the time difference (T) by obtaining this data.
Furthermore, selection processor 2302d identifies the first threshold value corresponding to the time difference (T), using the threshold value data (S204). Then, selection processor 2302d determines whether or not the sound volume ratio calculated is greater than or equal to the first threshold value (S205a).
When the sound volume ratio is greater than or equal to the first threshold value (“Yes” in S205a), selection processor 2302d selects the reflected sound as a reflected sound to be generated (S206). Specifically, in this case, selection processor 2302d selects that reproducer 2303 is to reproduce an output signal that is based on the audio signal that indicates the reflected sound and was created by analyzer 2301.
When the sound volume ratio is less than the first threshold value (“No” in S205a), selection processor 2302d skips selecting the reflected sound as a reflected sound to be generated (S207). Specifically, in this case, selection processor 2302d selects that reproducer 2303 is not to reproduce the output signal that is based on the audio signal that indicates the reflected sound and was created by analyzer 2301, thereby determining that the reflected sound is a reflected sound that is not to be generated, i.e., a reflected sound to be culled.
Subsequently, selection processor 2302d determines whether there are any unspecified reflected sounds (S208). That is, selection processor 2302d determines whether any of the plurality of audio signals created by analyzer 2301 have not undergone selection processing. If there are any unspecified reflected sounds (“Yes” in S208), selection processor 2302d repeats the above-described processing (S201 to S207, S210, and S220). If there are no unspecified reflected sounds (“No” in S208), selection processor 2302d ends the processing.
The selection processing is thus performed. In step S206, when selection processor 2302d selects that reproducer 2303 is to reproduce the output signal that is based on the audio signal indicating the reflected sound, selector 2302 outputs the audio signal to reproducer 2303.
Once more, a description is provided with reference to FIG. 30. Reproducer 2303 obtains the audio signal output from selector 2302 and outputs an output signal that is based on the audio signal (S103a). Here, reproducer 2303 synthesizes and outputs the audio signal that indicates the direct sound and was obtained by obtainer 2302a, and the sound signal generated (the audio signal indicating the reflected sound).
Thus, in step S205a, when the sound volume ratio is greater than or equal to the first threshold value, that is, when selection processor 2302d selects that reproducer 2303 is to reproduce an output signal that is based on the audio signal indicating the reflected sound, reproducer 2303 outputs the output signal that is based on that audio signal.
It should be noted that when the sound volume ratio is less than the first threshold value in step S205a, that is, when selection processor 2302d does not select that reproducer 2303 is to reproduce an output signal that is based on the audio signal indicating the reflected sound, reproducer 2303 skips outputting the output signal that is based on that audio signal. In such cases, reproducer 2303 does not output an output signal that is based on the audio signal, thereby reducing the amount of computation and the computational load.
As described above, the audio signal processing method according to the present embodiment is an audio signal processing method executed by an audio signal processing device (renderer 2300), and includes an obtaining step, a first calculating step, a second calculating step, a selection processing step, and a reproducing step.
In the obtaining step, an audio signal that includes attribute information identifying the attribute of the audio signal is obtained. The attribute includes information indicating an indirect sound (e.g., a reflected sound). In the first calculating step, a first sound volume is calculated. The first sound volume is based on the sound volume of the indirect sound (the sound volume of the reflected sound) at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present. The calculating of the first sound volume is based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained. In the second calculating step, a second sound volume is calculated. The second sound volume is based on a sound volume of a direct sound, associated with the indirect sound, at a time at which the direct sound arrives at the listening position. In the selection processing step, whether to output an output signal that is based on the audio signal obtained is selected, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives. In the reproducing step, the output signal is output, when outputting the output signal is selected.
Whether an output signal based on the audio signal indicating the indirect sound is output is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference (T). In other words, whether to output the output signal that is based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Furthermore, the first sound volume is calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the first sound volume calculated is used; and the time difference (T). That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Furthermore, in the audio signal processing method according to the present embodiment, in the second calculating step, the second sound volume is calculated based on the second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
The second sound volume is thus calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the second sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into further consideration.
Furthermore, in the audio signal processing method according to the present embodiment, the frequency characteristic indicating the auditory sensitivity is the frequency characteristic indicating the sound volume sensitivity of the listener, and the frequency characteristic indicating the sound volume sensitivity is the A-weighting characteristic.
This makes it possible to realize an audio signal processing method that enables using the frequency characteristic indicating the sound volume sensitivity of the listener as the frequency characteristic indicating the auditory sensitivity, and enables using the A-weighting characteristic as the frequency characteristic indicating the sound volume sensitivity.
Embodiment 3
Embodiment 3 is described below. The description below is centered on the points of difference from Embodiments 1 and 2, and descriptions of points in common are omitted or simplified.
In Embodiment 2, the output signal was selected to be output when the sound volume ratio between the second sound volume and the first sound volume was greater than or equal to the first threshold value. Embodiment 3 is similar to Embodiment 2 in the respect that the output signal is selected to be output when the sound volume ratio (L) is greater than or equal to the first threshold value. However, the method for determining the first threshold value in Embodiment 3 differs from that in Embodiment 2.
[Configuration of Renderer]
First, the configuration of renderer 3300 according to the present embodiment is described. FIG. 39 is a block diagram illustrating a configuration example of renderer 3300 according to the present embodiment.
Renderer 3300 includes analyzer 3301, selector 3302, and reproducer 2303.
Analyzer 3301 differs from analyzer 2301 according to Embodiment 2 in that analyzer 3301 calculates, e.g., a value related to the sound volume ratio (L) between a direct sound and a reflected sound at the listening position.
That is, analyzer 3301 detects direct sound and reflected sound that may be generated in the sound space, and when such direct sound and reflected sound are detected, analyzer 3301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on spatial information and sound data.
In the present embodiment, a direct sound and a reflected sound that is a sound resulting from the direct sound being reflected by a reflector (e.g., a non-sound-emitting object) are detected by analyzer 3301. The direct sound is also a sound associated with the reflected sound. Analyzer 3301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound associated with the reflected sound.
Furthermore, as in Embodiments 1 and 2, analyzer 3301 may calculate, for each of the direct sound and reflected sound, values related to: the path until arriving at the listening position; the time period taken until arrival; the sound volume at arrival; and the like. Analyzer 3301 then calculates values representing information indicating the relationship between the direct sound and the reflected sound, such as, for example, a value related to the time difference (T) between the direct sound and the reflected sound (the time difference (T) between when the direct sound arrives and when the reflected sound arrives), a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, and the like. The method by which analyzer 3301 calculates these items of information is as described in Embodiment 1.
It should be noted that as described in Embodiment 2, the sound volume of the reflected sound (the sound volume of the indirect sound) is the reflected sound arrival time sound volume (Ir), and the sound volume of the direct sound is the direct sound arrival time sound volume (Id). In other words, the sound volume ratio (L), calculated by analyzer 3301, between the direct sound and the reflected sound at the listening position is the sound volume ratio (L) between the sound volume of the direct sound and the sound volume of the reflected sound.
It should be noted that analyzer 3301 may be able to perform all of the processing or some of the processing performed by analyzer 1301 according to Embodiment 1.
Selector 3302 has obtainer 3302a and selection processor 3302d. Selector 3302 may be able to perform all of the processing or some of the processing performed by selector 1302 according to Embodiment 1.
Obtainer 3302a obtains the audio signal indicating the reflected sound and the audio signal indicating the direct sound associated with the reflected sound, both created by analyzer 3301. Furthermore, obtainer 3302a obtains, e.g., a value related to the time difference (T) between the direct sound and the reflected sound calculated by analyzer 3301 and a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, each calculated by analyzer 3301.
Selection processor 3302d selects whether reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound, based on: the sound volume ratio (L) and the time difference (T) obtained by obtainer 3302a; and the reflection coefficient characteristic amount.
For example, selection processor 3302d performs selection processing as described in step S102 of FIG. 8 in Embodiment 1 and step S102a of FIG. 30 in Embodiment 2. Selection processor 3302d selects that reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound, when the sound volume ratio (L) calculated by analyzer 3301 is greater than or equal to the first threshold value determined according to the time difference (T) between the direct sound and the reflected sound.
Furthermore, the time difference (T) is the time difference (T) between: the direct sound associated with the reflected sound indicated by the audio signal obtained; and the reflected sound indicated by the audio signal obtained, and is the time difference (T) between the direct sound and the reflected sound calculated by analyzer 3301. As described in Embodiment 1, the time difference (T) between a direct sound and a reflected sound is, for example, the time difference between the direct sound arrival time period (arrival time) and the reflected sound arrival time period (arrival time), but is not limited thereto.
The first threshold value is a value determined according to the time difference (T) between the direct sound associated with the reflected sound and the reflected sound; in other words, the threshold value is a value dependent on the time difference (T) and is the value indicated by the threshold value data of Embodiment 1.
Moreover, the reflection coefficient characteristic amount is a characteristic amount determined based on the reflection coefficient of the reflector (e.g., a non-sound-emitting object). The reflection coefficient characteristic amount is a characteristic amount indicating the degree of flatness of the frequency characteristic of the reflection coefficient, such as that illustrated in FIG. 32. More specifically, the reflection coefficient characteristic amount is a characteristic amount indicating the degree of flatness of the frequency spectrum of the gain characteristic indicated by the reflection coefficient. For the wall surface having a hard surface with numerous irregularities and the wall surface having a hard surface with few irregularities, each shown in FIG. 32, the degree of flatness of the frequency characteristic of the reflection coefficient is high, and the reflection coefficient characteristic amount is large. For the wall surface having a soft surface illustrated in FIG. 32, the degree of flatness of the frequency characteristic of the reflection coefficient is low, and the reflection coefficient characteristic amount is small.
Furthermore, the reflection coefficient characteristic amount can also be described as a characteristic amount indicating the degree of variation in the attenuation amount, as indicated by the reflection coefficients shown in FIG. 32.
The reflection coefficient characteristic amount may be obtained by obtainer 3302a via a communication line or the like. Furthermore, the reflection coefficient characteristic amount may be stored in the memory included in analyzer 3301, and obtainer 3302a may obtain the reflection coefficient characteristic amount stored in the memory.
When selection processor 3302d performs the selection processing, the reflection coefficient characteristic amount is used as follows.
In the present embodiment, the first threshold value is changed according to the magnitude of the reflection coefficient characteristic amount. For example, when the reflection coefficient characteristic amount is large, the first threshold value is changed to be higher, and when the reflection coefficient characteristic amount is small, the first threshold value is changed to be lower.
FIG. 40 is a diagram illustrating the impact of the reflection coefficient characteristic amount on the first threshold value, according to the present embodiment. FIG. 40 illustrates an example of the echo detection limit threshold value (first threshold value), similar to FIG. 28. When the reflection coefficient characteristic amount is large, the first threshold value becomes higher, shifting such that the first threshold value increases across the entire range of the horizontal axis shown in FIG. 40, for example. When the reflection coefficient characteristic amount is small, the first threshold value becomes lower, shifting such that the first threshold value becomes lower across the entire range of the horizontal axis, as shown in FIG. 40, for example.
Here, the precedence effect and the reflection coefficient characteristic amount are examined.
The technique described in Embodiment 1 and the like, described above, utilizes the precedence effect. The precedence effect is said to occur when the frequency spectrum of a leading sound (for example, a direct sound) approximates that of a lagging sound (for example, a reflected sound). In other words, when the frequency spectrum of the leading sound does not approximate that of the lagging sound, it is considered that the precedence effect will not occur.
Incidentally, since reflected sound is sound resulting from a direct sound being reflected by a reflector (e.g., a non-sound-emitting object), the frequency spectrum of the lagging sound (reflected sound) depends on the reflection coefficient of the reflector, more specifically, on the reflection coefficient characteristic amount. Therefore, whether the frequency spectrum of the direct sound approximates the frequency spectrum of the reflected sound varies depending on the reflection coefficient characteristic amount.
For example, when the reflection coefficient of the reflector corresponds to the reflection coefficients of the wall surface having a hard surface and numerous irregularities and the wall surface having a hard surface and few irregularities, each shown in FIG. 32, the reflection coefficient characteristic amount is large. Consequently, the frequency spectrum of the leading sound approximates the frequency spectrum of the lagging sound, making the precedence effect more likely to occur. In this case, the first threshold value is changed to be higher, that is, selection is performed such that the output signal that is based on the audio signal is less likely to be output. This is because the auditory value of reflected sound decreases as the precedence effect is more likely to occur.
Furthermore, when the reflection coefficient of the reflector is that of the wall surface having a soft surface shown in FIG. 32, the reflection coefficient characteristic amount is small. Consequently, the frequency spectrum of the leading sound does not approximate the frequency spectrum of the lagging sound, making the precedence effect less likely to occur. In this case, the first threshold value is changed to be lower, that is, selection is performed such that the output signal based on the audio signal is more likely to be output.
Selecting whether to output the output signal based on the reflection coefficient characteristic amount is thus equivalent to selecting whether to output the output signal while taking the precedence effect into consideration. Since the precedence effect is an example of the auditory sensitivity characteristic, the audio signal processing method according to the present embodiment is able to appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
As described above, selection processor 3302d selects whether reproducer 2303 is to output an output signal based on the obtained audio signal indicating the reflected sound, based on: the sound volume ratio (L) and the time difference (T) obtained by obtainer 3302a; and the reflection coefficient characteristic amount.
When selection processor 3302d selects that reproducer 2303 is to output the output signal, selector 3302 outputs the audio signal obtained to reproducer 2303.
Reproducer 2303 obtains the audio signal output from selector 3302 and outputs an output signal that is based on the audio signal obtained.
Hereinafter, an example of the operation of the audio signal processing method performed by the audio signal processing device according to the present embodiment (more specifically, renderer 3300) is described.
[Operation Example of Renderer]
FIG. 41 is a flowchart illustrating an operation example of the audio signal processing device according to the present embodiment. FIG. 41 illustrates the processing performed mainly by renderer 3300 included in the audio signal processing device according to the present embodiment.
First, analyzer 3301 performs analysis processing to analyze the input signal (S101b). More specifically, analyzer 3301 analyzes the input signal to detect direct sound and reflected sound that may be generated in the sound space. When such direct sound and reflected sound are detected, analyzer 3301 creates audio signals including attribute information, i.e., an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. Analyzer 3301 causes the created audio signal to be stored in the memory of analyzer 3301. Furthermore, analyzer 3301 analyzes the input signal and calculates, for each of the direct sound and reflected sound, values related to, e.g., the path until arrival at the listening position, the time taken to arrive, and the sound volume at arrival; a value related to the time difference (T) between the direct sound and the reflected sound; and a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position.
That is, analyzer 3301 calculates the sound volume ratio (L), which is the ratio between the direct sound arrival time sound volume (Id) and the reflected sound arrival time sound volume (Ir), and the time difference (T) between the direct sound and the reflected sound. The sound volume ratio (L) is the sound volume ratio (L) between the direct sound and the reflected sound at the listening position. It should be noted that the methods shown in Embodiment 1 may be used as the methods for calculating the sound volume ratio (L) and the time difference (T).
It should be noted that unlike step S101a according to Embodiment 2, in step S101b according to the present embodiment, the sound volume ratio (L) is calculated.
Selector 3302 (more specifically, selection processor 3302d) performs the selection of reflected sounds (selection processing) (S102b). In other words, selector 3302 selects whether reproducer 2303 is to reproduce an output signal based on the audio signal that indicates the reflected sound and was created by analyzer 3301.
First, obtainer 3302a obtains an audio signal that includes attribute information and was created by analyzer 3301 and stored in the memory. Obtainer 3302a obtains, for example, an audio signal indicating a reflected sound and an audio signal indicating a direct sound associated with the reflected sound. Furthermore, obtainer 3302a obtains, e.g., a value related to the time difference (T) between the direct sound and the reflected sound calculated by analyzer 3301 and a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, each calculated by analyzer 3301. Furthermore, obtainer 3302a obtains the reflection coefficient characteristic amount stored in the memory of analyzer 3301.
Selection processor 3302d selects that reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound, when the sound volume ratio (L) obtained is greater than or equal to the first threshold value.
The first threshold value is a value determined based on the time difference (T) between the direct sound and the reflected sound, and determined based on the reflection coefficient characteristic amount obtained. That is, selection processor 3302d determines the first threshold value based on the time difference (T) between the direct sound and the reflected sound, and the reflection coefficient characteristic amount.
As described above, selection processor 3302d selects whether reproducer 2303 is to output an output signal that is based on the audio signal obtained, based on the sound volume ratio (L) between the sound volume of the reflected sound and the sound volume of the direct sound, the time difference (T) between the direct sound and the reflected sound, and the reflection coefficient characteristic amount.
The selection processing is thus performed. When selection processor 3302d selects that reproducer 2303 is to reproduce the output signal that is based on the audio signal indicating the reflected sound, selector 3302 outputs the audio signal to reproducer 2303.
Then, reproducer 2303 obtains the audio signal output from selector 3302 and outputs an output signal that is based on the audio signal (S103b). Here, for example, reproducer 2303 synthesizes and outputs the audio signal indicating the direct sound obtained by obtainer 3302a and the sound signal generated (the audio signal indicating the reflected sound).
As described above, the audio signal processing method according to the present embodiment is an audio signal processing method executed by an audio signal processing device, and includes an obtaining step, a selection processing step, and a reproducing step.
In the obtaining step, an audio signal including attribute information identifying an attribute of the audio signal is obtained. The attribute includes information indicating a reflected sound that is a sound resulting from a direct sound being reflected by a reflector. That is, this attribute includes information indicating reflected sound. In the selection processing step, whether to output an output signal that is based on the audio signal obtained is selected, based on: a sound volume ratio between: a sound volume of the reflected sound, indicated by the audio signal obtained, at a time at which the reflected sound arrives at a listening position that is a position at which a listener is present; and a sound volume of the direct sound at a time at which the direct sound arrives at the listening position; a time difference between when the direct sound arrives and when the reflected sound arrives; and a reflection coefficient characteristic amount determined based on a reflection coefficient of the reflector. In the reproducing step, the output signal is output, when outputting the output signal is selected.
Whether to output the output signal that is based on the audio signal indicating the reflected sound is thus selected based on the above-described sound volume ratio and the above-described time difference. In other words, whether to output the output signal that is based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Here, attention is directed to the precedence effect. The precedence effect is said to occur when the frequency spectrum of a leading sound (for example, a direct sound) approximates the frequency spectrum of a lagging sound (for example, a reflected sound). The frequency spectrum of reflected sound varies according to the reflection coefficient of the reflector. Therefore, whether the frequency spectrum of a direct sound approximates the frequency spectrum of a reflected sound varies in accordance with the reflection coefficient characteristic amount.
Selecting whether to output the output signal based on the reflection coefficient characteristic amount is thus equivalent to selecting whether to output the output signal while taking the precedence effect into consideration. Since the precedence effect is an example of an auditory sensitivity characteristic, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking the auditory sensitivity into consideration.
Furthermore, in the present embodiment, the reflection coefficient characteristic amount indicates the degree of flatness of the frequency characteristic of the reflection coefficient.
This makes it possible to realize an audio signal processing method that enables using the reflection coefficient characteristic amount indicating the degree of flatness in the frequency characteristics of the reflection coefficient.
Embodiment 4
Embodiment 4 is described below. The description below is centered on the points of difference from Embodiment 2, and descriptions of points in common are omitted or simplified.
In Embodiment 2, the output signal was selected to be output when the sound volume ratio between the second sound volume and the first sound volume was greater than or equal to the first threshold value. Embodiment 4 differs from Embodiment 2 in that whether to output the output signal is selected based on a predetermined sound volume corresponding to the sound volume of a sound at a time at which the sound arrives at the listener.
It should be noted that in Embodiment 2 and the like, indirect sound (reflected sound) and direct sound were distinguished. However, in the present embodiment, when distinguishing between indirect sound (reflected sound) and direct sound is unnecessary, both indirect sound (reflected sound) and direct sound may simply be referred to as “sound”.
[Configuration of Renderer]
First, the configuration of renderer 4300 according to the present embodiment is described. FIG. 42 is a block diagram illustrating a configuration example of renderer 4300 according to the present embodiment.
Renderer 4300 includes analyzer 4301, selector 4302, and reproducer 2303.
Analyzer 4301 detects sounds that may be generated in the sound space. When such sound is detected, analyzer 4301 creates an audio signal indicating that sound, based on spatial information and sound data.
More specifically, analyzer 4301 detects direct sound and reflected sound that may be generated in the sound space, and when such direct sound and reflected sound are detected, analyzer 4301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. In other words, the audio signal indicating the sound is either an audio signal indicating reflected sound or an audio signal indicating direct sound.
In the present embodiment, a direct sound and a reflected sound that is a sound resulting from the direct sound being reflected by a reflector (e.g., a non-sound-emitting object) are detected by analyzer 4301. Thus, analyzer 4301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound associated with the reflected sound.
Furthermore, analyzer 4301 may calculate values related to the path the sound takes until arriving at the listening position, the time the sound takes to arrive, the arrival time sound volume, and the like. That is, as in Embodiments 1 and 2, for each of the direct sound and reflected sound, values related to the following may be calculated: the path until arriving at the listening position; the time period taken until arrival; the sound volume at arrival; and the like.
It should be noted that in the present embodiment, analyzer 4301 may not calculate values representing information indicating the relationship between the direct sound and the reflected sound, such as, for example, a value related to the time difference (T) between the direct sound and the reflected sound (the time difference (T) between when the direct sound arrives and when the reflected sound arrives), a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, and the like.
The audio signal according to the present embodiment includes information indicating the sound volume of the sound indicated by the audio signal, at the time at which the sound arrives at the listening position. That is, similar to Embodiment 2, an audio signal whose attribute is information indicating reflected sound includes information indicating the reflected sound arrival time sound volume (Ir), as the sound volume of the reflected sound at the time at which the reflected sound arrives at the listening position. An audio signal whose attribute is information indicating direct sound includes information indicating the direct sound arrival time sound volume (Id), as the sound volume of the direct sound at the time at which the direct sound arrives at the listening position.
It should be noted that analyzer 4301 may be able to perform all of the processing or some of the processing performed by analyzer 1301 according to Embodiment 1.
Selector 4302 has obtainer 4302a, third calculator 4302e, and selection processor 4302d.
Obtainer 4302a obtains audio signals that indicate sounds and were created by analyzer 4301. That is, obtainer 4302a obtains an audio signal indicating a reflected sound and an audio signal indicating a direct sound associated with the reflected sound.
It should be noted that as described above, the audio signals indicating sounds include information indicating the sound volume at the time at which the sound represented by that audio signal arrives at the listening position. That is, the audio signal indicating the reflected sound includes information representing the sound volume of the reflected sound (the reflected sound arrival time sound volume (Ir)), and the audio signal indicating the direct sound includes information representing the sound volume of the direct sound (the direct sound arrival time sound volume (Id)).
Third calculator 4302e calculates a predetermined sound volume based on the sound volume of the sound indicated by the audio signal at a time at which the sound arrives at the listening position, based on the audio signal indicating the sound. More specifically, third calculator 4302e calculates the predetermined sound volume based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal obtained by obtainer 4302a; and the audio signal obtained.
In the present embodiment, the audio signal indicating the sound is either an audio signal indicating reflected sound or an audio signal indicating direct sound.
When the audio signal indicating the sound is an audio signal indicating reflected sound, third calculator 4302e performs the following processing to calculate the predetermined sound volume.
In this case, the gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal is the same as the gain characteristic of each predetermined frequency bandwidth related to the reflected sound, described in Embodiment 2. Similarly, the frequency characteristic indicating the auditory sensitivity is, for example, an A-weighted characteristic. Further, the predetermined correction characteristic is the first correction characteristic described in Embodiment 2. Then, third calculator 4302e calculates the first sound volume as the predetermined sound volume by the same method as that in Embodiment 2, based on the predetermined correction characteristic (first correction characteristic) and the audio signal indicating the reflected sound. In other words, in this case, the predetermined sound volume is the first sound volume.
Furthermore, when the audio signal indicating the sound is an audio signal indicating direct sound, third calculator 4302e performs the following processing to calculate the predetermined sound volume.
In this case, the gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal is the same as the gain characteristic of each predetermined frequency bandwidth related to the direct sound, described in Embodiment 2. Similarly, the frequency characteristic indicating the auditory sensitivity is, for example, an A-weighted characteristic. Further, the predetermined correction characteristic is the second correction characteristic described in Embodiment 2. Then, third calculator 4302e calculates the second sound volume as the predetermined sound volume by the same method as that in Embodiment 2, based on the predetermined correction characteristic (second correction characteristic) and the audio signal indicating the direct sound. In other words, in this case, the predetermined sound volume is the second sound volume.
Selection processor 4302d selects whether reproducer 2303 is to output an output signal that is based on the audio signal obtained, based on the predetermined sound volume calculated by third calculator 4302e.
Selection processor 4302d selects that reproducer 2303 is to output an output signal that is based on the audio signal obtained, when the predetermined sound volume calculated is greater than or equal to the second threshold value.
The second threshold value, unlike the first threshold value, is a value independent of the time difference (T) between the direct sound associated with the reflected sound, and the reflected sound, and is a fixed value. The second threshold value is a value related to the sound volume obtained; in other words, the second threshold value is a value related to the amplitude value. Furthermore, the second threshold value indicates the sound volume of the boundary demarcating whether a sound is perceivable to the listener, and is a threshold value for determining a sound having a lower sound volume than the threshold value to be a sound that is not to be reproduced.
FIG. 43 is a graph illustrating threshold value data indicating the second threshold value according to the present embodiment. For example, the second threshold value is −70 dB. Since the predetermined sound volume (the second sound volume) of the audio signal indicating the direct sound in FIG. 43 is greater than or equal to the second threshold value, the output signal that is based on the audio signal is selected to be output. Furthermore, since the predetermined sound volume (the first sound volume) of the audio signal indicating the reflected sound in FIG. 43 is less than the second threshold value, the output signal that is based on that audio signal is selected not to be output.
It should be noted that selector 4302 may be able to perform all of the processing or some of the processing performed by selector 1302 according to Embodiment 1.
Hereinafter, an example of the operation of the audio signal processing method performed by the audio signal processing device according to the present embodiment (more specifically, renderer 4300) is described.
[Operation Example of Renderer]
FIG. 44 is a flowchart illustrating an operation example of the audio signal processing device according to the present embodiment. FIG. 44 illustrates the processing performed mainly by renderer 4300 included in the audio signal processing device according to the present embodiment.
First, analyzer 4301 performs analysis processing to analyze the input signal (S101c). More specifically, analyzer 4301 analyzes the input signal to detect sounds (direct sound and reflected sound) that may be generated in the sound space. When such sounds are detected, analyzer 4301 creates audio signals indicating the sounds, more specifically an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. Analyzer 4301 causes the audio signal created to be stored in the memory of analyzer 4301.
Furthermore, analyzer 4301 analyzes the input signal to calculate, for each sound (the direct sound and the reflected sound), values related to the path until arrival at the listening position, the time taken until arrival, the arrival time sound volume, and the like.
It should be noted that in step S101c according to the present embodiment, it is not necessary to calculate values such as a value related to the time difference (T) between the direct sound and the reflected sound, a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, and the like.
Selector 4302 (more specifically, selection processor 4302d) performs the selection of sounds (selection processing) (S102c). In other words, selector 4302 selects whether reproducer 2303 is to reproduce an output signal that is based on the audio signal that indicates the sound and was created by analyzer 4301.
First, obtainer 4302a obtains the audio signals (the audio signal indicating the reflected sound and the audio signal indicating the direct sound) that include attribute information and were created by analyzer 4301 and stored in the memory.
Then, third calculator 4302e calculates the predetermined sound volume based on the predetermined correction characteristic obtained by correcting, using the frequency characteristic indicating the auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal obtained by obtainer 4302a; and the audio signal obtained.
When the audio signal indicating the sound is an audio signal indicating reflected sound, the predetermined sound volume is the first sound volume. Furthermore, when the audio signal indicating the sound is an audio signal indicating direct sound, the predetermined sound volume is the second sound volume.
Next, selection processor 4302d selects that reproducer 2303 is to output an output signal based on the audio signal obtained, when the predetermined sound volume calculated by third calculator 4302e is greater than or equal to the second threshold value.
Selector 4302 performs the selection processing as described above.
The selection processing is described in greater detail with reference to FIG. 45.
FIG. 45 is a flowchart illustrating an operation example of the selection processing according to the present embodiment.
First, selector 4302 specifies a sound detected by analyzer 4301 (S201c). In other words, obtainer 4302a of selector 4302 specifies the audio signal created by analyzer 4301 and stored in the memory, and obtains the audio signal specified.
Then, third calculator 4302e calculates the predetermined sound volume (the first sound volume or the second sound volume) (S230).
Next, selection processor 4302d determines whether or not the predetermined sound volume calculated is greater than or equal to the second threshold value (S205c).
When the predetermined sound volume is greater than or equal to the second threshold value (“Yes” in S205c), selection processor 4302d selects the sound indicated by the audio signal obtained as a sound to be generated (S206c). Specifically, in this case, selection processor 4302d selects that reproducer 2303 is to reproduce an output signal that is based on the audio signal that indicates the sound and was created by analyzer 4301.
When the predetermined sound volume is lower than the second threshold value (“No” in S205c), selection processor 4302d skips selecting the sound indicated by the audio signal obtained as a sound to be generated (S207c). Specifically, in this case, selection processor 4302d selects that reproducer 2303 is not to reproduce the output signal that is based on the audio signal that indicates the sound and was created by analyzer 4301, thereby determining that the sound is a sound that is not to be generated, i.e., a sound to be culled.
Subsequently, selection processor 4302d determines whether there are any unspecified sounds (S208c). That is, selection processor 4302d determines whether any of the plurality of audio signals created by analyzer 4301 have not undergone selection processing. If there are any unspecified sounds (“Yes” in S208c), selection processor 4302d repeats the above-described processing (S201c to S207c and S230). If there are no unspecified reflected sounds (“No” in S208c), selection processor 4302d ends the processing.
The selection processing is thus performed. In step S206c, when selection processor 4302d selects that reproducer 2303 is to reproduce the output signal based on the audio signal indicating the sound, selector 4302 outputs the audio signal to reproducer 2303.
Then, reproducer 2303 obtains the audio signal output from selector 4302 and outputs an output signal that is based on the audio signal (S103c).
As described above, the audio signal processing method according to the present embodiment is an audio signal processing method executed by an audio signal processing device, and includes an obtaining step, a third calculating step, a selection processing step, and a reproducing step.
In the obtaining step, an audio signal is obtained. In the third calculating step, a predetermined sound volume is calculated. The predetermined sound volume is based on a sound volume of a sound at a time at which the sound arrives at a listening position that is a position at which a listener is present. The sound is a sound indicated by the audio signal obtained. The calculating is based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound; and the audio signal obtained. In the selection processing step, whether to output an output signal that is based on the audio signal obtained is selected, based on the predetermined sound volume calculated. In the reproducing step, the output signal is output, when outputting the output signal is selected.
The predetermined sound volume is thus calculated while taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal that is based on the audio signal representing the sound is selected, based on the predetermined sound volume calculated. That is, whether to output the output signal that is based on the audio signal is appropriately selected while taking the auditory sensitivity into consideration. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
(Supplement)
Note that the aspects understood based on the present disclosure are not limited to the embodiment, and various changes may be performed.
For example, a process performed by a certain constituent element in the embodiment may be performed by another constituent element instead of the specific constituent element. Furthermore, the order of a plurality of processes may be changed, or a plurality of processes may be performed in parallel.
Moreover, ordinals such as first and second used for description may be interchanged, removed, or newly assigned as appropriate. These ordinals do not necessarily correspond to meaningful orders, and may be used to distinguish between elements.
Furthermore, for example, in comparisons between threshold values, “greater than or equal to” a threshold value and “greater than” a threshold value may be read interchangeably. Similarly, “less than or equal to” a threshold value and “less than” a threshold value may be read interchangeably. Moreover, for example, there may be cases in which the terms “time period” and “time” are read interchangeably.
Furthermore, in a process for selecting one or more sounds to be processed from a plurality of sounds, no sounds need be selected as a sound to be processed if no sounds that satisfy the conditions exist. In other words, a case in which no sounds to be processed are selected may be included in the process for selecting one or more sounds to be processed from a plurality of sounds.
Furthermore, at least one of a first element, a second element, or a third element can correspond to the first element, the second element, the third element, or any combination of these.
It should be noted that the information, described in Embodiment 2, indicating the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the gain characteristic of each predetermined frequency bandwidth related to the direct sound, and the frequency characteristic indicating the auditory sensitivity may be identified based on the spatial information included in the input information, for example. The metadata includes, for example, information representing the reflectance of structures that may reflect sound in the sound space, such as floors, walls, and ceilings, as well as information representing the reflectance of obstacle objects present in the sound space and the reflectance of their reflecting surfaces. Here, reflectance is defined as the ratio of energy or amplitude between reflected sound and incident sound, and is set for each frequency band of a sound. Furthermore, the reflectance may also be referred to as the reflection coefficient. For example, data such as that shown in FIG. 32 may be set in the metadata as information indicating the reflectance. FIG. 32 exemplifies the reflection coefficients for each of the three types of wall surfaces. For a non-sound-emitting object, one reflection coefficient may be set, or a plurality of reflection coefficients may be set. The gain of the equalizer applied to the reflected sound for each frequency band (i.e., the first correction characteristic) may be calculated based on the reflection coefficient set for the reflecting surface related to that reflected sound, described in the above example.
Furthermore, the gain characteristic of each predetermined frequency bandwidth related to a reflected sound (indirect sound), the gain characteristic of each predetermined frequency bandwidth related to a direct sound, and the frequency characteristic indicating the auditory sensitivity may be supplied using a communication line or the like, or may be stored in advance in the memory of analyzer 2301.
In addition, for example, in the embodiment, the case in which the aspects that are understood based on the present disclosure are implemented as an audio signal processing device, an encoding device, or a decoding device has been described. However, the aspects that are understood based on the present disclosure are not limited thereto, and may be implemented as software for executing the audio signal processing method, the encoding method, or the decoding method.
For example, a program for executing the above-described audio signal processing method, encoding method, or decoding method may be stored beforehand in ROM. Then, a CPU may operate according to this program.
Furthermore, a program for executing the above-described audio signal processing method, encoding method, or decoding method may be stored on a computer-readable recording medium. Then, a computer may record, in computer RAM, the program stored on the recording medium, and operate according to this program.
Moreover, each of the above-described constituent elements may be expressed typically as a large-scale integration (LSI), which is an integrated circuit (IC) having an input terminal and an output terminal. These may take the form of individual chips, or all or one or more constituent elements of the embodiment may be encapsulated in a single chip. Depending upon the level of integration, the LSI may be expressed as an IC, a system LSI, a super LSI, or an ultra LSI.
Furthermore, such IC is not limited to an LSI, and a dedicated circuit or a general-purpose processor may be used. Alternatively, a field programmable gate array (FPGA) that allows for programming after the manufacture of an LSI, or a reconfigurable processor that allows for reconfiguration of the connection and the setting of circuit cells inside an LSI may be employed. Furthermore, when a circuit integration technology that replaces LSIs comes along owing to advances in semiconductor technology or to a separate derivative technology, the constituent elements should naturally be integrated using that technology. The adaptation of biotechnology, and the like are also conceivable as possibilities.
Moreover, an FPGA, a CPU, or the like may, by means of wireless communication or wired communication, download all or a part of the software for executing the audio signal processing method, the encoding method, or the decoding method described in the present disclosure. Furthermore, all or a part of software for updating may be downloaded by means of wireless communication or wired communication. Moreover, an FPGA, a CPU, or the like may execute the digital signal processing described in the present disclosure by storing the downloaded software in memory and operating based on the stored software.
At this time, the machine that includes the FPGA, the CPU, or the like may be connected wirelessly or in a wired manner to a signal processing device, or may be connected to a signal processing server over a network. Accordingly, this machine and the signal processing device or the signal processing server may perform the audio signal processing method, the encoding method, or the decoding method described in the present disclosure.
For example, the audio signal processing device, the encoding device, or the decoding device in the present disclosure may include an FPGA, a CPU, or the like. Furthermore, the audio signal processing device, the encoding device, or the decoding device may include: an interface for acquiring, from an external source, the software for causing the FPGA, the CPU, or the like to operate; and memory for storing the acquired software. The FPGA, the CPU, or the like may perform the signal processing described in the present disclosure by operating based on the stored software.
A server may provide the software related to the acoustic processing, the encoding processing, or the decoding processing of the present disclosure. Furthermore, a terminal or a machine may operate as the audio signal processing device, the encoding device, or the decoding device described in the present disclosure by installing the software. Note that the terminal or the machine may install the software by connecting to a server over a network.
Furthermore, the software may be installed on the terminal or the machine by means of another device that is different from the terminal or the machine obtaining data for installing the software by connecting to a server over a network and providing the data for installing the software to the terminal or the machine. Note that VR software or AR software for causing a terminal or a machine to execute the audio signal processing method described by way of the embodiment may be an example of the software.
Note that in the foregoing embodiment, each constituent element may be configured from dedicated hardware, or may be implemented by executing a software program suitable for each constituent element. Each constituent element may be implemented by means of a program executor such as a CPU or a processor loading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
Thus, the device and the like according to one or more aspects have been described by way of the embodiment, but the aspects understood based on the present disclosure are not limited to the embodiment. The one or more aspects may thus include forms obtained by making various modifications to the above embodiments that can be conceived by those skilled in the art, as well as forms obtained by combining constituent elements in different variations, without materially departing from the spirit of the present disclosure.
INDUSTRIAL APPLICABILITY
The present disclosure includes aspects that can be applied to, for example, an audio signal processing device, an encoding device, a decoding device, or a terminal or equipment that includes any of these.
Publication Number: 20260230772
Publication Date: 2026-08-06
Assignee: Panasonic Intellectual Property Corporation Of America
Abstract
An audio signal processing method executed by an audio signal processing device includes: obtaining an audio signal having an attribute indicating an indirect sound; calculating a first sound volume based on a sound volume of the indirect sound when arriving at a listening position, using: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal; calculating a second sound volume based on a sound volume of a direct sound, associated with the indirect sound, when arriving at the listening position; selecting whether to output an output signal based on the audio signal, using: a sound volume ratio between the second and first sound volumes; and a time difference between the direct and indirect sounds; and outputting the output signal, when outputting the output signal is selected.
Claims
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
12.
13.
Description
CROSS REFERENCE TO RELATED APPLICATIONS
This is a continuation application of PCT International Application No. PCT/JP2024/035621 filed on Oct. 4, 2024, designating the United States of America, which is based on and claims priority of U.S. Provisional Patent Application No. 63/542,854 filed on Oct. 6, 2023. The entire disclosures of the above-identified applications, including the specifications, drawings and claims are incorporated herein by reference in their entirety.
FIELD
The present disclosure relates to an audio signal processing method and the like.
BACKGROUND
In recent years, the spread of products and services that utilize extended reality (ER) (may be also expressed as “XR”) including virtual reality (VR), augmented reality (AR), and mixed reality (MR) has advanced. Accompanying this, there has been growing demand for audio signal processing technologies that provide listeners with immersive audio that, in a virtual space or a real-world space, assigns acoustic effects that are generated in accordance with the environment of the space to sounds emitted from a virtual sound source.
Note that “listener” can also be expressed as “user”. Furthermore, Patent Literature (PTL) 1, PTL 2, PTL 3, and Non Patent Literature (NPL) 1 disclose techniques that relate to the audio signal processing method and the like of the present disclosure.
CITATION LIST
Patent Literature
PTL 1: Japanese Patent No. 6288100PTL 2: Japanese Unexamined Patent Application Publication No. 2019-22049PTL 3: WO Publication No. 2021/180938
Non Patent Literature
NPL 1: B. C. J. Moore, “An Introduction to the Psychology of Hearing”, Seishin Shobo, 1994 Apr. 20, Chapter 6: Space Perception, p. 225.
SUMMARY
Technical Problem
Incidentally, in the technique disclosed in PTL 1, it may be difficult to appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Accordingly, the present disclosure provides an audio signal processing method and the like that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Solution to Problem
An audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; calculating a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Furthermore, an audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating a reflected sound that is a sound resulting from a direct sound being reflected by a reflector; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between: a sound volume of the reflected sound at a time at which the reflected sound arrives at a listening position that is a position at which a listener is present; and a sound volume of the direct sound at a time at which the direct sound arrives at the listening position, the reflected sound being indicated by the audio signal obtained; a time difference between when the direct sound arrives and when the reflected sound arrives; and a reflection coefficient characteristic amount determined based on a reflection coefficient of the reflector; and outputting the output signal, when outputting the output signal is selected.
Furthermore, an audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal; calculating a predetermined sound volume, the predetermined sound volume being based on a sound volume of a sound at a time at which the sound arrives at a listening position that is a position at which a listener is present, the sound being a sound indicated by the audio signal obtained, the calculating being based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound; and the audio signal obtained; selecting whether to output an output signal that is based on the audio signal obtained, based on the predetermined sound volume calculated; and outputting the output signal, when outputting the output signal is selected.
Furthermore, an audio signal processing method according to one aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal; calculating a first sound volume, the first sound volume being based on a sound volume of an indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the audio signal for which the attribute includes information indicating the indirect sound, the attribute being identified by the attribute information included in the audio signal obtained; and gain information of the indirect sound; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound, the calculating of the second sound volume being based on: a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the direct sound; and gain information of the direct sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Furthermore, a recording medium according to one aspect of the present disclosure is a non-transitory computer-readable recording medium having recorded thereon a computer program for causing a computer to execute the audio signal processing method described above.
Furthermore, an audio signal processing device according to one aspect of the present disclosure is an audio signal processing device including: an obtainer that obtains an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; a first calculator that calculates a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the first calculator calculating the first sound volume based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; a second calculator that calculates a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; a selection processor that selects whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and a reproducer that outputs the output signal, when outputting the output signal is selected.
Note that these comprehensive or specific aspects may be implemented as a system, a device, a method, an integrated circuit, a computer program, or a non-transitory computer-readable recording medium such as a CD-ROM, or may be implemented as any combination of a system, a device, a method, an integrated circuit, a computer program, and a recording medium.
Advantageous Effects
The audio signal processing method and the like according to the one aspect of the present disclosure make it possible to appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
BRIEF DESCRIPTION OF DRAWINGS
These and other advantages and features will become apparent from the following description thereof taken in conjunction with the accompanying Drawings, by way of non-limiting examples of embodiments disclosed herein.
FIG. 1 is a diagram illustrating one example of a direct sound and reflected sounds generated in a sound space.
FIG. 2 is a diagram illustrating an example of a three-dimensional sound reproduction system according to Embodiment 1.
FIG. 3A is a block diagram illustrating a configuration example of an encoding device according to Embodiment 1.
FIG. 3B is a block diagram illustrating a configuration example of a decoding device according to Embodiment 1.
FIG. 3C is a block diagram illustrating another configuration example of an encoding device according to Embodiment 1.
FIG. 3D is a block diagram illustrating another configuration example of a decoding device according to Embodiment 1.
FIG. 4A is a block diagram illustrating a configuration example of a decoder according to Embodiment 1.
FIG. 4B is a block diagram illustrating another configuration example of a decoder according to Embodiment 1.
FIG. 5 is a diagram illustrating an example of a physical configuration of an audio signal processing device according to Embodiment 1.
FIG. 6 is a diagram illustrating an example of a physical configuration of an encoding device according to Embodiment 1.
FIG. 7 is a block diagram illustrating a configuration example of a renderer according to Embodiment 1.
FIG. 8 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 1.
FIG. 9 is a diagram illustrating a comparatively distant positional relationship between a listener and an obstacle object.
FIG. 10 is a diagram illustrating a comparatively close positional relationship between a listener and an obstacle object.
FIG. 11 is a diagram illustrating relationships between time differences between direct sounds and reflected sounds, and threshold values.
FIG. 12A is a diagram illustrating a part of an example of a method for setting threshold value data.
FIG. 12B is a diagram illustrating a part of an example of a method for setting threshold value data.
FIG. 12C is a diagram illustrating a part of an example of a method for setting threshold value data.
FIG. 13 is a diagram illustrating an example of a threshold value setting method.
FIG. 14 is a flowchart illustrating an example of selection processing.
FIG. 15 is a diagram illustrating relationships between directions of direct sounds, directions of reflected sounds, time differences, and threshold values.
FIG. 16 is a diagram illustrating relationships between angular differences, time differences, and threshold values.
FIG. 17 is a block diagram illustrating another configuration example of a renderer.
FIG. 18 is a flowchart illustrating another example of selection processing.
FIG. 19 is a flowchart illustrating yet another example of selection processing.
FIG. 20 is a flowchart illustrating a first variation of operations of an audio signal processing device according to Embodiment 1.
FIG. 21 is a flowchart illustrating a second variation of operations of an audio signal processing device according to Embodiment 1.
FIG. 22 is a diagram illustrating an arrangement example of an avatar, a sound source object, and an obstacle object.
FIG. 23 is a flowchart illustrating yet another example of selection processing.
FIG. 24 is a block diagram illustrating a configuration example for a renderer to perform pipeline processing.
FIG. 25 is a diagram illustrating transmission and diffraction of a sound.
FIG. 26 is a diagram illustrating an example of the positional relationship between a listener and an obstacle object, according to Embodiment 1.
FIG. 27 is a diagram illustrating another example of the positional relationship between a listener and an obstacle object, according to Embodiment 1.
FIG. 28 is an example of an echo detection limit threshold value according to Embodiment 1.
FIG. 29 is a block diagram illustrating a configuration example of a renderer according to Embodiment 2.
FIG. 30 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 2.
FIG. 31 is a flowchart illustrating an operation example of the selection processing according to Embodiment 2.
FIG. 32 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth related to a reflected sound according to Embodiment 2.
FIG. 33A is a diagram illustrating a table showing a frequency characteristic indicating auditory sensitivity according to Embodiment 2.
FIG. 33B is a diagram illustrating the frequency characteristic indicating the auditory sensitivity according to Embodiment 2.
FIG. 34 is a diagram illustrating the gain characteristic related to the reflected sound according to Embodiment 2, the frequency characteristic indicating the auditory sensitivity (A-weighting characteristic), and a first correction characteristic.
FIG. 35 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth related to a direct sound according to Embodiment 2.
FIG. 36 is a diagram illustrating the gain characteristic related to the direct sound, the frequency characteristic indicating auditory sensitivity (A-weighting characteristic), and a second correction characteristic, according to Embodiment 2.
FIG. 37 is a diagram illustrating the inverse characteristic of an equal loudness contour, which is another, first example of the frequency characteristic indicating the sound volume sensitivity of the listener according to Embodiment 2.
FIG. 38 is a diagram illustrating the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position, which is another, second example of the frequency characteristic indicating the sound volume sensitivity of the listener according to Embodiment 2.
FIG. 39 is a block diagram illustrating a configuration example of a renderer according to Embodiment 3.
FIG. 40 is a diagram illustrating the impact of a reflection coefficient characteristic amount on the first threshold value, according to Embodiment 3.
FIG. 41 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 3.
FIG. 42 is a block diagram illustrating a configuration example of a renderer according to Embodiment 4.
FIG. 43 is a graph illustrating threshold value data indicating a second threshold value according to Embodiment 4.
FIG. 44 is a flowchart illustrating an operation example of an audio signal processing device according to Embodiment 4.
FIG. 45 is a flowchart illustrating an operation example of the selection processing according to Embodiment 4.
DESCRIPTION OF EMBODIMENTS
(Underlying Knowledge Forming Basis of the Present Disclosure)
To date, investigations have been made into audio signal processing technologies that provide listeners with immersive audio by, in a virtual space or a real-world space, assigning acoustic effects that are generated in accordance with the environment of the space to sounds emitted from a virtual sound source.
PTL 1 discloses such an audio signal processing technique. More specifically, PTL 1 discloses a technique for detecting the importance of audio signals (voice signals) and not outputting audio signals for which the detected importance is low. By thus not outputting audio signals for which the importance is low, there are expectations for the audio signal processing technique to appropriately reduce the amount of computation and the computational load.
Incidentally, in a sound space (a virtual space or a real-world space), reflected sound is sometimes important.
FIG. 1 is a diagram illustrating one example of a direct sound and reflected sounds generated in a sound space. In acoustic processing in which characteristics of a virtual space are expressed by a sound, it is effective to reproduce not only direct sounds, but also reflected sounds in order to express the size of the space, the material of the walls, and the like, as well as to allow for accurately grasping the location of the sound source (the positioning of the sound image).
For example, when a sound is heard in a rectangular parallelepiped room such as that in FIG. 1, six primary reflected sounds, corresponding to the six walls, are generated for one sound source. Reproducing these reflected sounds provides a clue for appropriate understanding of the space and the sound image. Furthermore, for each reflected sound, a secondary reflected sound is generated by a surface other than the reflection surface that generated that reflected sound. These reflected sounds are also effective sensory clues.
However, even when consideration is given no further than to secondary reflection, one direct sound and 36 (6+6×5) reflected sounds are generated for one sound source. Thus, 37 sound rays are generated, and processing these sound rays requires a significant amount of computation.
Furthermore, in applied products in recent years for which metaverses are imagined, such as virtual meetings, virtual shopping, virtual concerts, and the like, a plurality of sound sources are present out of necessity, whereby an even greater amount of computation is required.
Moreover, the listener hearing the sounds in a virtual space uses headphones or VR goggles. In order to provide three-dimensional sound to such a listener, binaural processing that assigns a sound pressure ratio and a phase difference between the two ears and reproduces the direction of arrival and distance sensation of the sounds is performed on each sound ray. Thus, if an attempt were made to reproduce every reflected sound that is generated, the amount of computation would become immense.
On the other hand, in light of convenience, a small storage battery is sometimes used as the battery for the VR goggles worn by the listener who experiences the virtual space. Lessening the computational load resulting from the above-described processing makes it possible to further extend the life of the storage battery. To this end, the number of sound rays, which are emitted on a scale of hundreds, is desirably reduced, within a scope at which grasping the space and the positioning of the sounds is not harmed.
Furthermore, in a system that reproduces acoustics, a degree of freedom such as 6DoF (6 degrees of freedom) or the like may be allowed with respect to the position (in other words, the listening position, which is the position where the listener is present) and orientation of the listener. In this case, the positional relationship between the listener, the sound sources, and the objects that reflect sounds cannot be fixed until the time of reproduction (the time of rendering). For this reason, the reflected sounds as well cannot be fixed until the time of reproduction. Thus, it is difficult to determine the reflected sounds to be processed beforehand.
Therefore, during reproduction, appropriately selecting and outputting (reproducing) one or more reflected sounds, from a plurality of reflected sounds that are generated in a sound space, that are to be processed or are not to be processed is useful in appropriately reducing the amount of computation and the computational load.
It should be noted that controlling whether to select a sound corresponds to determining whether to select the sound, and more specifically corresponds to determining whether to select and output (reproduce) the sound. Furthermore, selecting a sound may be selecting the sound as a sound to be processed, or may be selecting the sound as a sound that is not to be processed.
Incidentally, in PTL 1, the importance of the audio signal, and more specifically the importance of the direct sound that the audio signal represents, is detected, but the importance of reflected sound is not considered. Therefore, when indirect sounds such as reflected sounds are generated, as illustrated in FIG. 1, the amount of computation and the computational load increase; in other words, it may be difficult to appropriately reduce the amount of computation and the computational load.
Furthermore, in conventional techniques including the technique disclosed in PTL 1, whether to output an audio signal is selected without taking the auditory sensitivity of the listener into consideration. When such an audio signal is output and the listener hears the sound indicated by that audio signal, the listener hears a sound that differs from his/her own auditory perception, causing a sense of incongruence.
Therefore, there is a need for an audio signal processing method and the like that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration, in a sound space.
Accordingly, an audio signal processing method according to a first aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; calculating a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Whether to output the output signal based on the audio signal indicating the indirect sound is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Furthermore, the first sound volume is calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the first sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
An audio signal processing method according to a second aspect of the present disclosure is the audio signal processing method according to the first aspect, wherein in the calculating of the second sound volume, the second sound volume is calculated based on a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the direct sound.
The second sound volume is thus calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the second sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into further consideration.
An audio signal processing method according to a third aspect of the present disclosure is the audio signal processing method according to the first or second aspect, wherein the frequency characteristic indicating the auditory sensitivity is a frequency characteristic indicating sound volume sensitivity of the listener.
This makes it possible to realize an audio signal processing method that enables using the frequency characteristic indicating the sound volume sensitivity of the listener as the frequency characteristic indicating the auditory sensitivity.
An audio signal processing method according to a fourth aspect of the present disclosure is the audio signal processing method according to the third aspect, wherein the frequency characteristic indicating the sound volume sensitivity is a frequency characteristic based on an inverse of an equal loudness contour.
This makes it possible to realize an audio signal processing method that enables using the frequency characteristic based on the inverse of the equal loudness contour as the frequency characteristic indicating the sound volume sensitivity.
An audio signal processing method according to a fifth aspect of the present disclosure is the audio signal processing method according to the third aspect, wherein the frequency characteristic indicating the sound volume sensitivity is an A-weighting characteristic.
This makes it possible to realize an audio signal processing method that enables using the A-weighting characteristic as the frequency characteristic indicating the sound volume sensitivity.
An audio signal processing method according to a sixth aspect of the present disclosure is the audio signal processing method according to the third aspect, wherein the frequency characteristic indicating the sound volume sensitivity is an inverse characteristic of a frequency characteristic of a minimum audible angle of a sound source position.
This makes it possible to realize an audio signal processing method that enables using the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position as the frequency characteristic indicating the sound volume sensitivity.
An audio signal processing method according to a seventh aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating a reflected sound that is a sound resulting from a direct sound being reflected by a reflector; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between: a sound volume of the reflected sound at a time at which the reflected sound arrives at a listening position that is a position at which a listener is present; and a sound volume of the direct sound at a time at which the direct sound arrives at the listening position, the reflected sound being indicated by the audio signal obtained; a time difference between when the direct sound arrives and when the reflected sound arrives; and a reflection coefficient characteristic amount determined based on a reflection coefficient of the reflector; and outputting the output signal, when outputting the output signal is selected.
Whether to output the output signal that is based on the audio signal indicating the reflected sound is thus selected based on the above-described sound volume ratio and the above-described time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Here, attention is directed to the precedence effect. The precedence effect is said to occur when the frequency spectrum of a leading sound (for example, a direct sound) approximates the frequency spectrum of a lagging sound (for example, a reflected sound). The frequency spectrum of reflected sound varies according to the reflection coefficient of the reflector. Therefore, whether the frequency spectrum of a direct sound approximates the frequency spectrum of a reflected sound varies in accordance with the reflection coefficient characteristic amount.
For example, when the reflection coefficient characteristic amount has a certain value, the frequency spectrum of the direct sound and the frequency spectrum of the reflected sound approximate each other, making it more likely for the precedence effect to occur. In this case, selection may be performed such that the output signal is less likely to be output. Furthermore, for example, when the reflection coefficient characteristic amount has another certain value, the frequency spectrum of the direct sound and the frequency spectrum of the reflected sound do not approximate each other, making it less likely for the precedence effect to occur. In this case, selection may be performed such that the output signal is more likely to be output.
That is, selecting whether to output the output signal based on the reflection coefficient characteristic amount is equivalent to selecting whether to output the output signal while taking the precedence effect into consideration. Since the precedence effect is an example of an auditory sensitivity characteristic, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
An audio signal processing method according to an eighth aspect of the present disclosure is the audio signal processing method according to the seventh aspect, wherein the reflection coefficient characteristic amount is a characteristic amount indicating a degree of flatness of a frequency characteristic of a reflection coefficient.
This makes it possible to realize an audio signal processing method that enables using, as the reflection coefficient characteristic amount, the characteristic amount indicating the degree of flatness of the frequency characteristic of the reflection coefficient.
An audio signal processing method according to a ninth aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal; calculating a predetermined sound volume, the predetermined sound volume being based on a sound volume of a sound at a time at which the sound arrives at a listening position that is a position at which a listener is present, the sound being a sound indicated by the audio signal obtained, the calculating being based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound; and the audio signal obtained; selecting whether to output an output signal that is based on the audio signal obtained, based on the predetermined sound volume calculated; and outputting the output signal, when outputting the output signal is selected.
The predetermined sound volume is thus calculated while taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal that is based on the audio signal representing the sound is selected, based on the predetermined sound volume calculated. That is, whether to output the output signal that is based on the audio signal is appropriately selected while taking the auditory sensitivity into consideration. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
An audio signal processing method according to a tenth aspect of the present disclosure is an audio signal processing method executed by an audio signal processing device, the audio signal processing method including: obtaining an audio signal including attribute information identifying an attribute of the audio signal; calculating a first sound volume, the first sound volume being based on a sound volume of an indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the calculating of the first sound volume being based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the audio signal for which the attribute includes information indicating the indirect sound, the attribute being identified by the attribute information included in the audio signal obtained; and gain information of the indirect sound; calculating a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound, the calculating of the second sound volume being based on: a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth of the direct sound; and gain information of the direct sound; selecting whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and outputting the output signal, when outputting the output signal is selected.
Whether to output the output signal based on the audio signal indicating the indirect sound is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
An audio signal processing method according to an eleventh aspect of the present disclosure is the audio signal processing method according to the tenth aspect, wherein the gain characteristic of each predetermined frequency bandwidth related to the indirect sound is stored as the attribute information related to the indirect sound, the gain information related to the indirect sound is stored as the attribute information of the indirect sound, the gain characteristic of each predetermined frequency bandwidth related to the direct sound is stored as the attribute information related to the direct sound, and the gain information related to the direct sound is stored as the attribute information of the direct sound.
This makes it possible to realize an audio signal processing method in which various information is stored as the attribute information.
A recording medium according to a twelfth aspect of the present disclosure is a non-transitory computer-readable recording medium having recorded thereon a computer program for causing a computer to execute the audio signal processing method according to any one of the first to eleventh aspects.
This makes it possible for a computer to execute the above-described audio signal processing method, according to the computer program.
An audio signal processing device according to a thirteenth aspect of the present disclosure is an audio signal processing device including: an obtainer that obtains an audio signal including attribute information identifying an attribute of the audio signal, the attribute including information indicating an indirect sound; a first calculator that calculates a first sound volume, the first sound volume being based on a sound volume of the indirect sound at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present, the first calculator calculating the first sound volume based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained; a second calculator that calculates a second sound volume, the second sound volume being based on a sound volume of a direct sound at a time at which the direct sound arrives at the listening position, the direct sound being associated with the indirect sound; a selection processor that selects whether to output an output signal that is based on the audio signal obtained, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives; and a reproducer that outputs the output signal, when outputting the output signal is selected.
Whether to output the output signal based on the audio signal indicating the indirect sound is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference. In other words, whether to output the output signal based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Furthermore, the first sound volume is calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the first sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing device that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Embodiment 1
(Example of Three-Dimensional Sound Reproduction System)
FIG. 2 is a diagram illustrating an example of three-dimensional sound reproduction system 1000. Specifically, FIG. 2 illustrates three-dimensional sound reproduction system 1000, which is an example of a system to which acoustic processing or decoding processing of the present disclosure can be applied. Three-dimensional sound is also expressed as immersive audio. Three-dimensional sound reproduction system 1000 includes audio signal processing device 1001 and audio presentation device 1002.
Audio signal processing device 1001, which is also expressed as an acoustic processing device, applies acoustic processing to an audio signal emitted from a virtual sound source and generates an acoustic-processed audio signal to be presented to the listener. The audio signal is not limited to voices, and is acceptable as long as it is an audible sound. Acoustic processing is, for example, signal processing applied to an audio signal in order to reproduce one or more effects that a sound receives between when the sound is emitted from a sound source and when the sound arrives at the listener.
Audio signal processing device 1001 performs acoustic processing based on spatial information that describes the main factors for bringing about the above-described effects. Spatial information encompasses, for example: information that indicates the location of a sound source, a listener, and objects in the vicinity; information that indicates the shape of a space; parameters regarding sound propagation; and the like. Audio signal processing device 1001 is, for example, a PC (personal computer), a smartphone, a tablet, a game console, or the like.
An acoustic-processed signal is presented from audio presentation device 1002 to the listener. Audio presentation device 1002 is connected to audio signal processing device 1001 via wireless or wired communication. The acoustic-processed audio signal generated by audio signal processing device 1001 is transmitted to audio presentation device 1002 via wireless or wired communication.
When audio presentation device 1002 includes a plurality of devices such as, for example, a device for the right ear and a device for the left ear, or the like, the plurality of devices present sound in synchronization by means of communication between the plurality of devices or communication between each of the plurality of devices and audio signal processing device 1001. Audio presentation device 1002 is, for example, headphones, earphones, or a head-mounted display worn on the head of the listener, surround speakers including a plurality of fixed speakers, or the like.
Note that three-dimensional sound reproduction system 1000 may be used in combination with an image presentation device or a stereoscopic image presentation device that visually provides an ER experience that includes AR/VR. For example, a space handled by spatial information is a virtual space in which the positions of sound sources, the listener, and objects in the space are virtual positions of virtual sound sources, a virtual listener, and virtual objects in a virtual space. The space can also be expressed as a sound space. Furthermore, the spatial information can also be expressed as sound space information.
Furthermore, FIG. 2 illustrates a system configuration example in which audio signal processing device 1001 and audio presentation device 1002 are separate devices, but three-dimensional sound reproduction system 1000 to which the acoustic processing method (the audio signal processing method) or the decoding method of the present disclosure can be applied is not limited to the configuration in FIG. 2. For example, audio signal processing device 1001 may be included in audio presentation device 1002, and audio presentation device 1002 may perform both acoustic processing and sound presentation.
Furthermore, audio signal processing device 1001 and audio presentation device 1002 may, in a shared manner, perform the acoustic processing described in the present disclosure. Furthermore, a server connected to audio signal processing device 1001 or audio presentation device 1002 over a network may perform a part or all of the acoustic processing described in the present disclosure.
Furthermore, audio signal processing device 1001 may perform the acoustic processing by decoding a bitstream that has been generated by encoding at least a part of data of the audio signal and the spatial information used in the acoustic processing. Thus, audio signal processing device 1001 may be expressed as a decoding device.
(Example of Encoding Device)
FIG. 3A is a block diagram illustrating a configuration example of encoding device 1100. Specifically, FIG. 3A illustrates the configuration of encoding device 1100, which is an example of the encoding device of the present disclosure.
Input data 1101 is data to be encoded, and includes spatial information and/or an audio signal to be inputted into encoder 1102. Details regarding the spatial information will be described later.
Encoder 1102 encodes input data 1101 to generate encoded data 1103. Encoded data 1103 is, for example, a bitstream generated by means of encoding processing.
Memory 1104 stores encoded data 1103. Memory 1104 may be, for example, a hard disk or an SSD (solid-state drive), or may be another type of memory.
Note that in the above description, a bitstream generated by means of encoding processing was given as an example of encoded data 1103 stored in memory 1104, but encoded data 1103 may be data other than a bitstream. For example, encoding device 1100 may store, in memory 1104, converted data generated by converting the bitstream into a predetermined data format. The converted data may be, for example, a file or multiplexed stream that corresponds to one or more bitstreams.
Here, the file is a file having a file format of, for example, ISO base media file format (ISOBMFF) or the like. Furthermore, encoded data 1103 may be in the form of a plurality of packets generated by splitting the above-described bitstream or file.
For example, the bitstream generated by encoder 1102 may be converted to data that is different from the bitstream. In this case, encoding device 1100 may include a converter, not illustrated, and the converter may perform conversion processing, or conversion processing may be performed by a central processing unit (CPU) that is an example of a processor, described later.
(Example of Decoding Device)
FIG. 3B is a block diagram illustrating a configuration example of decoding device 1110. Specifically, FIG. 3B illustrates the configuration of decoding device 1110, which is an example of the decoding device of the present disclosure.
Memory 1114 stores, for example, the same data as encoded data 1103 generated by encoding device 1100. The stored data is read from memory 1114 and inputted into decoder 1112 as input data 1113. Input data 1113 is, for example, a bitstream that is to be decoded. Memory 1114 may be, for example, a hard disk or an SSD, or may be another type of memory.
Note that decoding device 1110 may not directly input, to decoder 1112, data read from memory 1114 as input data 1113, and may instead convert the data read and then input the converted data to decoder 1112 as input data 1113. The data before conversion may be, for example, multiplexed data that includes one or more bitstreams. Here, the multiplexed data may be, for example, a file having a file format such as ISOBMFF or the like.
Furthermore, the data before conversion may be a plurality of packets generated by splitting the above-described bitstream or file. Data that is different from the bitstream may be read from memory 1114 and then converted into a bitstream. In this case, decoding device 1110 may include a converter, not illustrated, and the converter may perform conversion processing, or conversion processing may be performed by a CPU that is an example of a processor, described later.
Decoder 1112 decodes input data 1113 to generate audio signal 1111 that indicates audio to be presented to the listener.
(Other Example of Encoding Device)
FIG. 3C is a block diagram illustrating a configuration example of another encoding device. Specifically, FIG. 3C illustrates the configuration of encoding device 1120, which is another example of the encoding device of the present disclosure. In FIG. 3C, constituent elements that are the same as the constituent elements in FIG. 3A have been given the same reference signs as in FIG. 3A, and description of these constituent elements is omitted.
Encoding device 1100 stores encoded data 1103 in memory 1104. On the other hand, encoding device 1120 is different from encoding device 1100 in the respect that encoding device 1120 includes transmitter 1121 that transmits encoded data 1103 externally.
Transmitter 1121 transmits, to a different device or server, transmission signal 1122 that is generated based on data converted from encoded data 1103 or encoded data 1103 to a different file format. The data used in generating transmission signal 1122 is, for example, the bitstream, multiplexed data, file, or packet described in relation to encoding device 1100.
(Other Example of Decoding Device)
FIG. 3D is a block diagram illustrating another configuration example of a decoding device. Specifically, FIG. 3D illustrates the configuration of decoding device 1130, which is another example of the decoding device of the present disclosure. In FIG. 3D, constituent elements that are the same as the constituent elements in FIG. 3B have been given the same reference signs as in FIG. 3B, and description of these constituent elements is omitted.
Decoding device 1110 reads input data 1113 from memory 1114. On the other hand, decoding device 1130 is different from decoding device 1110 in the respect that decoding device 1130 includes receiver 1131, which receives input data 1113 from an external source.
Receiver 1131 receives reception signal 1132 to obtain reception data, and outputs input data 1113 to be inputted into decoder 1112. The reception data may be the same as input data 1113 inputted into decoder 1112, or may be data in a data format that is different from that of input data 1113.
When the data format of the reception data is different from the data format of input data 1113, receiver 1131 may convert the reception data into input data 1113. Alternatively, a converter or a CPU, each not illustrated, of decoding device 1130 may convert the reception data into input data 1113. The reception data is, for example, the bitstream, multiplexed data, file, or packet described in relation to encoding device 1120.
(Example of Decoder)
FIG. 4A is a block diagram illustrating a configuration example of decoder 1200. Specifically, FIG. 4A illustrates the configuration of decoder 1200, which is an example of decoder 1112 in FIG. 3B and FIG. 3D.
Input data 1113 is an encoded bitstream, and includes encoded audio data that is an audio signal that has been encoded, and metadata used in acoustic processing.
Spatial information manager 1201 obtains the metadata included in input data 1113 and analyzes the metadata. The metadata includes information that describes the main factors that act on the sounds arranged in the sound space. Spatial information manager 1201 manages the spatial information that is obtained by analyzing the metadata and is used in the acoustic processing, and provides the spatial information to renderer 1203.
Note that in the present disclosure, the information used in the acoustic processing is expressed as spatial information, but another expression may be used. For example, the information used in the acoustic processing may be expressed as sound space information, or may be expressed as scene information. Furthermore, when the information used in the acoustic processing changes over time, the spatial information inputted into renderer 1203 may be information expressed as a spatial state, a sound space state, a scene state, or the like.
Furthermore, the spatial information may be managed for each sound space or for each scene. For example, when each of a plurality of mutually differing rooms is expressed as a virtual space, the plurality of rooms may be managed as a plurality of scenes that mutually differ. Furthermore, spatial information may be managed such that even the same room is managed as a different scene in accordance with the expressed state.
Thus, a plurality of items of spatial information may be managed with respect to a plurality of sound spaces or a plurality of scenes. In management of a plurality of items of spatial information, an identifier that identifies each item of the plurality of items of spatial information may be assigned to the spatial information.
The spatial information data may be included in a bitstream that is an example of input data 1113. Alternatively, the bitstream may include an identifier of the spatial information, and the spatial information data may be obtained from an information source other than the bitstream. Specifically, when the bitstream includes only the identifier of the spatial information, in the rendering, the spatial information data stored in the memory inside the device or in an external server may be obtained as input data 1113, using the identifier of the spatial information.
Note that the information managed by spatial information manager 1201 is not limited to information included in the bitstream. For example, input data 1113 may include, as data not included in the bitstream, data that indicates the characteristics and structure of a space obtained from a VR or AR software application or server.
Furthermore, input data 1113 may include data that indicates the characteristics, position, and/or the like of the listener or an object. Moreover, input data 1113 may include information on the position of the listener, obtained using a sensor included in a terminal including a decoding device (1110, 1130), or may include information that indicates the position of the terminal, estimated based on information obtained using the sensor.
In other words, spatial information manager 1201 may communicate with an external system or server to obtain spatial information and the position of the listener (in other words, the listening position). Furthermore, spatial information manager 1201 may obtain clock synchronization information from an external system and perform processing to perform synchronization with a clock in renderer 1203.
Note that the space in the above description may be a virtually formed space, i.e., a VR space, or may be a real-world space or a virtual space that corresponds to a real-world space, i.e., an AR space or an MR space. Furthermore, the virtual space may be expressed as a sound field or a sound space. Moreover, the information indicating position in the above description may be information on coordinates or the like that indicate a position in a space, may be information that indicates a relative position with respect to a predetermined reference position, or may be information that indicates movement or acceleration of a position in a space.
Audio data decoder 1202 decodes encoded audio data included in input data 1113 to obtain an audio signal.
The encoded audio data obtained by three-dimensional sound reproduction system 1000 is, for example, a bitstream encoded in a predetermined format such as MPEG-H 3D Audio (ISO/IEC 23008-3). Note that MPEG-H 3D Audio is merely an example of an encoding method that can be used when generating the encoded audio data included in the bitstream. The encoded audio data may be a bitstream encoded by another encoding method.
For example, the encoding method may be a lossy codec such as MPEG-1 Audio Layer III (MP3), Advanced Audio Coding (AAC), Windows Media Audio (WMA), Audio Codec 3 (AC3), Vorbis, or the like. Alternatively, the encoding method may be a lossless codec such as Apple Lossless Audio Codec (ALAC), Free Lossless Audio Codec (FLAC), or the like.
Alternatively, any encoding method other than the above-described may be used. For example, PCM data may be a type of the encoded audio data. In this case, when, for example, the quantization bit rate of the PCM data is N, the decoding processing may be processing in which the N-bit binary number is converted into a numerical format (for example, floating-point format) that can be processed by renderer 1203.
Renderer 1203 obtains the audio signal and the spatial information, applies acoustic processing to the audio signal using the spatial information, and outputs an acoustic-processed audio signal (audio signal 1111).
Before starting the rendering, spatial information manager 1201 reads the metadata of the input signal, detects rendering items such as objects and sounds specified by the spatial information, and transmits the rendering items to renderer 1203. After the start of rendering, spatial information manager 1201 grasps the change over time of the spatial information and the position of the listener, and updates and manages the spatial information. Then, the updated spatial information is transmitted to renderer 1203.
Renderer 1203 generates and outputs an acoustic processing-added audio signal based on the audio signal included in input data 1113 and the spatial information received from spatial information manager 1201.
Update processing of the spatial information and output processing of the acoustic processing-added audio signal may be performed in the same thread. Furthermore, processing may be allocated to spatial information manager 1201 and renderer 1203 in mutually independent threads. When spatial information manager 1201 and renderer 1203 perform the update processing of the spatial information and the output processing of the acoustic processing-added audio signal in different threads, the thread activation frequency may be set separately, or processing may be performed in parallel.
When spatial information manager 1201 and renderer 1203 perform processing in independent threads that are different, renderer 1203 can be preferentially allotted computation resources. This makes it possible to safely execute sound output processing for which even a slight delay is unallowable, i.e., for which a popping noise would be generated with a delay of even one sample (0.02 msec).
At this time, allotment of computation resources to spatial information manager 1201 is limited. However, since compared to the output processing of the audio signal, updating of the spatial information is processing that is infrequent (for example, processing such as updating the orientation of a listener's face), it is not necessary for the output processing of the audio signal to be performed instantaneously. Thus, even when the allotment of computation resources is limited, the acoustic quality is not greatly impacted.
The updating of the spatial information may be periodically performed with each elapse of a preset time period or term, or may be performed when a preset condition is satisfied. Furthermore, the updating of the spatial information may be performed manually by the listener or a sound space manager, or may be performed by being triggered by a change to an external system.
For example, the spatial information may be updated when a controller is being operated by the listener and the position of the listener's avatar is instantaneously warped or time is instantaneously progressed or reversed. Alternatively, the spatial information may be updated when an operation to suddenly change the field environment is performed by the manager of the virtual space. In these cases, the thread for updating the spatial information managed by spatial information manager 1201 may be activated as one-time interrupt processing in addition to the periodic activation.
The role of the information update thread that performs update processing of the spatial information is, for example: processing to update, based on the position or orientation of the VR goggles worn by the listener, the position or orientation of the listener's avatar positioned in the virtual space; updating of the positions of objects that have moved within the virtual space; and the like. This is handled within a processing thread that operates at a relatively low frequency of around several tens of Hz. Processing that reflects the characteristics of direct sound may also be performed in such a low-frequency processing thread. This is because the characteristics of a direct sound vary less frequently than audio processing frames for audio output occur. Rather, by doing so, the computational load of such processing can be made relatively low, and the risk of pulsive noise can be avoided, since updating information at an unduly high frequency would generate the risk of pulsive noise occurring.
FIG. 4B is a block diagram illustrating another configuration example of a decoder. Specifically, FIG. 4B illustrates the configuration of decoder 1210, which is another example of decoder 1112 in FIG. 3B and FIG. 3D.
FIG. 4B is different from FIG. 4A in the respect that input data 1113 includes not encoded audio data, but an unencoded audio signal. Input data 1113 includes an audio signal and a bitstream including metadata.
Spatial information manager 1211 is the same as spatial information manager 1201 in FIG. 4A; therefore, description thereof has been omitted.
Renderer 1213 is the same as renderer 1203 in FIG. 4A; therefore, description thereof has been omitted.
Note that decoders 1112, 1200, and 1210 may be expressed as the acoustic processor that performs the acoustic processing. Furthermore, decoding devices 1110 and 1130 may be audio signal processing device 1001, or may be expressed as the acoustic processing device.
(Physical Configuration of Audio Signal Processing Device)
FIG. 5 is a diagram illustrating an example of a physical configuration of audio signal processing device 1001. Note that audio signal processing device 1001 in FIG. 5 may be decoding device 1110 in FIG. 3B or decoding device 1130 in FIG. 3D. A plurality of the constituent elements illustrated in FIG. 3B or FIG. 3D may be implemented by a plurality of the constituent elements illustrated in FIG. 5. Furthermore, a part of the configuration described here may be included in audio presentation device 1002.
Audio signal processing device 1001 in FIG. 5 includes processor 1402, memory 1404, communication interface (I/F) 1403, sensor 1405, and loudspeaker 1401.
Processor 1402 is, for example, a CPU, a digital signal processor (DSP), or a graphics processing unit (GPU). The acoustic processing or the decoding processing of the present disclosure may be performed by the CPU, the DSP, or the GPU executing a program stored in memory 1404. Furthermore, processor 1402 is, for example, a circuit that performs information processing. Processor 1402 may be a dedicated circuit that performs signal processing on audio signals, including the acoustic processing of the present disclosure.
Memory 1404 includes, for example, random access memory (RAM) or read-only memory (ROM). Memory 1404 may include, for example, magnetic storage media, exemplified by a hard disk, or semiconductor memory, exemplified by an SSD. Furthermore, memory 1404 may be an internal memory incorporated into the CPU or GPU. Moreover, spatial information managed by spatial information manager 1201 and 1211 and/or the like may be stored in memory 1404. Furthermore, threshold value data, described later, may be stored.
Communication I/F 1403 is, for example, a communication module that supports a communication method such as Bluetooth (registered trademark) or WiGig (registered trademark). Audio signal processing device 1001 communicates with other communication devices via communication I/F 1403, and obtains a bitstream to be decoded. The obtained bitstream is, for example, stored in memory 1404.
Communication I/F 1403 includes, for example, a signal processing circuit that supports the communication method, and an antenna. The communication method is not limited to Bluetooth (registered trademark) or WiGig (registered trademark), and may be Long Term Evolution (LTE), New Radio (NR), Wi-Fi (registered trademark), or the like.
The communication method is not limited to the wireless communication methods described above, and may be a wired communication method such as Ethernet (registered trademark), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI) (registered trademark), or the like.
Sensor 1405 performs sensing to estimate the position or orientation of the listener. Specifically, sensor 1405 estimates the position and/or orientation of the listener based on one or more detection results of one or more of the position, orientation, movement, velocity, angular velocity, acceleration, or the like of a part or all of the listener's body, and generates position/orientation information indicating the position and/or orientation of the listener.
Note that a device outside of audio signal processing device 1001 may include sensor 1405. The part of the body may be the listener's head or the like. The position/orientation information may be information indicating the position and/or orientation of the listener in real-world space, or may be information indicating the displacement of the position and/or orientation of the listener with respect to the position and/or orientation of the listener at a predetermined time point. Furthermore, the position/orientation information may be information indicating a position and/or orientation relative to three-dimensional sound reproduction system 1000 or an external device including sensor 1405.
Sensor 1405 may be, for example, an imaging device such as a camera or a distance measuring device such as a laser imaging detection and ranging (LIDAR) distance measuring device. Sensor 1405 may capture an image of the movement of the listener's head and detect the movement of the listener's head by processing the captured image. Furthermore, a device that performs position estimation using radio waves in any given frequency band such as millimeter waves may be used as sensor 1405.
Furthermore, audio signal processing device 1001 may obtain position information via communication I/F 1403 from an external device including sensor 1405. In this case, audio signal processing device 1001 need not include sensor 1405. Here, the external device refers to, for example, audio presentation device 1002 described in FIG. 2, or a stereoscopic image reproduction device worn on the listener's head. In this case, sensor 1405 is configured as a combination of various sensors, such as a gyro sensor and an acceleration sensor, for example.
As the speed of the movement of the listener's head, sensor 1405 may detect, for example, the angular speed of rotation about at least one of three mutually orthogonal axes in the sound space as the axis of rotation or the acceleration of displacement in at least one of the three axes as the direction of displacement.
As the amount of the movement of the listener's head, sensor 1405 may detect, for example, the amount of rotation about at least one of three mutually orthogonal axes in the sound space as the axis of rotation or the amount of displacement in at least one of the three axes as the direction of displacement. Specifically, sensor 1405 detects 6DoF positions (x, y, z) and angles (yaw, pitch, roll) as the position of the listener. Sensor 1405 is configured as a combination of various sensors used for detecting movement, such as a gyro sensor and an acceleration sensor.
Note that sensor 1405 may implemented by, e.g., a camera or a Global Positioning System (GPS) receiver for detecting the position of the listener. Position information obtained by performing self-position estimation, by using LIDAR or the like as sensor 1405, may be used. For example, when three-dimensional sound reproduction system 1000 is implemented by a smartphone, sensor 1405 is included in the smartphone.
Furthermore, sensor 1405 may include a temperature sensor such as a thermocouple that detects the temperature of audio signal processing device 1001. Moreover, sensor 1405 may include, for example, a sensor that detects the remaining level of a battery included in audio signal processing device 1001 or a battery connected to audio signal processing device 1001.
Loudspeaker 1401 includes, for example, a diaphragm, a driving mechanism such as a magnet or a voice coil, and an amplifier, and presents the acoustic-processed audio signal as sound to the listener. Loudspeaker 1401 operates the driving mechanism according to the audio signal (more specifically, a waveform signal indicating the waveform of the sound) amplified via the amplifier, and vibrates the diaphragm by means of the driving mechanism. In this way, the diaphragm vibrating according to the audio signal generates sound waves, which propagate through the air and are transmitted to the listener's ears, allowing the listener to perceive the sound.
Note that although here, an example in which audio signal processing device 1001 includes loudspeaker 1401 and presents the acoustic-processed audio signal via loudspeaker 1401 was given, the means for providing the audio signal is not limited to this configuration.
For example, the acoustic-processed audio signal may be outputted to external audio presentation device 1002 connected via a communication module. The communication performed by the communication module may be wired or wireless. As another example, audio signal processing device 1001 may include a terminal that outputs an analog audio signal, and may present the audio signal from earphones or the like by connecting the earphone cable to the terminal.
In this case, audio presentation device 1002 may be headphones, earphones, a head-mounted display, neck speakers, wearable speakers, or the like, each worn on the listener's head or a part of the listener's body. Alternatively, audio presentation device 1002 may be surround speakers configured with a plurality of fixed speakers, or the like. Audio presentation device 1002 may reproduce the audio signal.
(Physical Configuration of Encoding Device)
FIG. 6 is a diagram illustrating an example of a physical configuration of encoding device 1500. Encoding device 1500 in FIG. 6 may be encoding device 1100 in FIG. 3A or encoding device 1120 in FIG. 3C, or a plurality of the constituent elements illustrated in FIG. 3A or FIG. 3C may be implemented by a plurality of the constituent elements illustrated in FIG. 6.
Encoding device 1500 in FIG. 6 includes processor 1501, memory 1503, and communication I/F 1502.
Processor 1501 is, for example, a CPU, a DSP, or a GPU. The encoding processing of the present disclosure may be performed by the CPU, the DSP, or the GPU executing a program stored in memory 1503. Furthermore, processor 1501 is, for example, a circuit that performs information processing. Processor 1501 may be a dedicated circuit that performs signal processing on audio signals, including the encoding processing of the present disclosure.
Memory 1503 includes, for example, RAM or ROM. Memory 1503 may include, for example, magnetic storage media, exemplified by a hard disk, or semiconductor memory, exemplified by an SSD. Furthermore, memory 1503 may be an internal memory incorporated into the CPU or GPU.
Communication I/F 1502 is, for example, a communication module that supports a communication method such as Bluetooth (registered trademark) or WiGig (registered trademark). For example, encoding device 1500 communicates with other communication devices via communication I/F 1502, and transmits an encoded bitstream.
Communication I/F 1502 includes, for example, a signal processing circuit that supports the communication method, and an antenna. The communication method is not limited to Bluetooth (registered trademark) or WiGig (registered trademark), and may be LTE, NR, Wi-Fi (registered trademark), or the like. The communication method is not limited to wireless communication methods. The communication method may be a wired communication method such as Ethernet (registered trademark), USB, HDMI (registered trademark), or the like.
The communication module includes, for example, a signal processing circuit that supports the communication method, and an antenna. In the above example, Bluetooth (registered trademark) and WiGig (registered trademark) were given as examples of the communication method, but a communication method such as Long Term Evolution (LTE), New Radio (NR), Wi-Fi (registered trademark), or the like may be supported. Furthermore, the communication I/F may be not the wireless communication methods described above, but a wired communication method such as Ethernet (registered trademark), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI) (registered trademark), or the like.
[Configuration of Renderer]
FIG. 7 is a block diagram illustrating a configuration example of renderer 1300. Specifically, FIG. 7 illustrates an example of the detailed configuration of renderer 1300, which corresponds to renderers 1203 and 1213 in FIG. 4A and FIG. 4B.
Renderer 1300 includes analyzer 1301, selector 1302, and reproducer 1303, and adds acoustic processing to sound data included in the input signal and outputs the sound data.
The input signal includes, for example, spatial information, sensor information, and sound data. The input signal may include a bitstream that includes sound data and metadata (control information), and in this case, the spatial information may be included in the metadata.
The spatial information is information related to the sound space (three-dimensional sound field) created by three-dimensional sound reproduction system 1000, and includes information about objects included in the sound space and information about the listener. The objects include sound source objects that emit sound and serve as sound sources, and non-sound-emitting objects that do not emit sound. The sound source objects may be expressed as simply sound sources.
The non-sound-emitting object serves as an obstacle object that reflects sound emitted by the sound source object, but a sound source object may also serve as an obstacle object that reflects sound emitted by another sound source object. The obstacle object may also be expressed as a reflection object.
Information assigned in common to both sound source objects and non-sound-emitting objects includes position information, geometry information, and the attenuation rate of sound volume when the object reflects sound.
The position information is represented by coordinate values of three axes, for example, the X-axis, the Y-axis, and the Z-axis of Euclidean space, but it does not necessarily have to be three-dimensional information. For example, the position information may be two-dimensional information represented by coordinate values of the two axes of the X-axis and the Y-axis. The position information of the object is defined by a representative position of the shape expressed by a mesh or voxel.
The geometry information may include information about the material of the surface.
The attenuation rate may be expressed as a real number greater than or equal to 0 and less than or equal to 1, or may be expressed as a negative decibel value. Since sound volume does not increase from reflection in real-world space, the attenuation rate is set to a negative decibel value. However, for example, to create an eerie atmosphere in a non-realistic space, an attenuation rate greater than or equal to 1, that is, a positive decibel value, may be intentionally set.
Furthermore, the attenuation rate may be set such that each frequency band included in a plurality of frequency bands has a different value, or values may be independently set for each frequency band. Furthermore, when the attenuation rate is set for each type of material of an object surface, the value of the corresponding attenuation rate may be used based on information about the surface material.
Furthermore, the spatial information may include, for example, information indicating whether the object belongs to an animate thing or information indicating whether the object is a mobile body. When the object is a mobile body, the position indicated by the position information may move over time. In this case, information on the changed position or the amount of change is transmitted to renderer 1300.
Information related to the sound source object includes, in addition to information assigned in common to both sound source objects and non-sound-emitting objects, sound data and information necessary for radiating the sound data into the sound space. The sound data is data representing sound perceived by the listener, and indicates information such as the frequency and intensity of the sound.
The sound data is typically a PCM signal, but may also be data compressed using an encoding method such as MP3. In this case, since the signal needs to be decoded at least before arriving at reproducer 1303, renderer 1300 may include a decoder (not illustrated). Alternatively, the signal may be decoded by audio data decoder 1202.
One item of sound data may be set for one sound source object, or a plurality of items of sound data may be set. Furthermore, identification information for identifying each item of sound data may be assigned to the sound data, or information related to the sound source object may include identification information for the sound data.
As information necessary for radiating sound data into the sound space, for example, information on a reference sound volume that is used as a standard in reproducing the sound data, information indicating a characteristic of sound data, information related to the position of the sound source object, information related to the orientation of the sound source object (in other words, information related to the directivity of the sound emitted by the sound source object), and the like may be included.
The information on the reference sound volume may be, for example, the root mean square value of the amplitude value of the sound data at the sound source position at the time of radiating the sound data into the sound space, and may be expressed as a floating-point decibel (dB) value.
For example, when the reference sound volume is 0 dB, the information on the reference sound volume may indicate that the sound is to be radiated into the sound space from the position indicated by the information related to the position of the sound source object, at the same sound volume as the signal level indicated by the sound data, without increase or decrease. Furthermore, when the reference sound volume is-6 dB, this may indicate that the sound is to be radiated into the sound space from the position indicated by the information related to the position of the sound source object at approximately half the sound volume of the signal level indicated by the sound data.
The information on the reference sound volume may be assigned to each item of sound data, or may be assigned collectively to a plurality of items of sound data.
The information indicating a characteristic of sound data is, for example, information on the sound volume of a sound source, and may be information indicating the chronological variation in the sound volume of a sound source.
For example, when the sound space is a virtual conference room and the sound source is a speaker, the sound volume transitions intermittently over short periods of time. In other words sound portions and silent portions occur alternately. Furthermore, when the sound space is a concert hall and the sound source is a performer, the sound volume is maintained over a certain duration of time. Moreover, when the sound space is a battlefield and the sound source is an explosive, the sound volume of the explosion sound becomes large for only an instant and then continues to be silent or in a quiet state thereafter.
In this way, the sound volume information on the sound source may include not only information on the magnitude of sound but also information on the transition of the sound magnitude. Such information may be used as the information indicating a characteristic of the sound data.
The information on the transition may be expressed by data indicating frequency characteristics in chronological order. The information on the transition may be expressed by data indicating the duration of a sound interval. The information on the transition may be expressed by data indicating the chronological order of durations of sound intervals and durations of silent intervals. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration for which the amplitude of the sound signal can be considered stationary (can be considered approximately constant) and the amplitude value of said signal during that duration.
The information on the transition may be expressed by data of a duration during which the frequency characteristics of the sound signal can be considered stationary. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration during which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during that duration. The information on the transition may be expressed in the format of, for example, data indicating the general shape of a spectrogram.
Furthermore, the sound volume that is used as the standard for the above-mentioned frequency characteristics may be used as the reference sound volume. The information on the reference sound volume and the information indicating a characteristic of the sound data may be used for calculation processing of the sound volume of direct sound or reflected sound to be perceived by the listener, and/or may be used for selection processing for selecting whether to make the listener perceive the sound. Other examples of and usage methods for the information indicating a characteristic of sound data will be described later.
It should be noted that the reflected sound according to the present embodiment is an example of indirect sound. Indirect sound may be reflected sound, diffracted sound, or the like. In the present embodiment, reflected sound, which is an example of indirect sound, is used for description, but the same processing is performed even if indirect sound other than reflected sound is used.
Information regarding the orientation of a sound source object (orientation information) is typically expressed in terms of yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted, and the orientation information of a sound source object may be expressed in terms of azimuth (yaw) and elevation (pitch). The orientation information of a sound source object may change over time, and when changed, the orientation information is transmitted to renderer 1300.
Information related to the listener is information regarding the position and orientation of the listener in the sound space. The information regarding the position (position information) is represented by the position on the X-, Y-, and Z-axes of Euclidean space, but need not necessarily be three-dimensional information and may be two-dimensional information. Information regarding the orientation of the listener (orientation information) is typically expressed in terms of yaw, pitch, and roll. Alternatively, the rotation of roll may be omitted, and the listener orientation information may be expressed in terms of azimuth (yaw) and elevation (pitch).
The position information and orientation information regarding a listener may change over time, and when changed, the position information and orientation information are transmitted to renderer 1300.
The sensor information is information that includes, e.g., the rotation amount or displacement amount detected by sensor 1405 worn by the listener, and the position and orientation of the listener. The sensor information is transmitted to renderer 1300, and renderer 1300 updates the information on the position and orientation of the listener based on the sensor information. The sensor information may include position information obtained by performing self-localization estimation by a mobile terminal using GPS, a camera, or LIDAR, for example.
Furthermore, information obtained not from sensor 1405, but from an external source through a communication module, may also be detected as sensor information. Information indicating the temperature of audio signal processing device 1001, and information indicating the remaining level of the battery may be obtained from sensor 1405. Moreover, computational resources (CPU capability, memory resources, PC performance, and the like) of audio signal processing device 1001 or audio presentation device 1002 may be obtained in real time.
Analyzer 1301 analyzes an audio signal included in the input signal and spatial information received from spatial information managers 1201 and 1211 to calculate the information required for generating direct sounds and reflected sounds with reproducer 1303, and the information required for selecting whether to generate reflected sounds.
The information required for generating direct sounds and reflected sounds is, for example, for each of direct sounds and reflected sounds, values related to the path until arriving at the listening position, the time period taken until arrival, the sound volume at the arrival time, and the like. The values related to, e.g., the path until arriving at the listening position, the time period taken until arrival, and the sound volume at the arrival time are, for example, values representing the path until arriving at the listening position, the time period taken until arrival, and the sound volume at the arrival time, respectively.
The information required for selecting a reflected sound to be output is information indicating the relationship between the direct sound and the reflected sound, and is, for example, a value regarding a time difference between the direct sound and the reflected sound, a value regarding a sound volume ratio of the reflected sound to the direct sound at the listening position, and/or the like. The value related to the time difference between a direct sound and a reflected sound and the value related to the sound volume ratio between a direct sound and a reflected sound at the listening position are, for example, a value representing the time difference between the direct sound and the reflected sound and a value representing the sound volume ratio of the reflected sound to the direct sound at the listening position, respectively.
Note that it goes without saying that when the sound volume is expressed in units of decibels on a logarithmic scale (when the sound volume is expressed in the decibel domain), the sound volume ratio between the two signals is expressed as a decibel value difference. Specifically, the sound volume ratio between the two signals may be the difference when the amplitude value of each signal is expressed in the decibel domain. That value may be calculated based on, e.g., an energy value, a power value, or the like. Furthermore, this difference can be referred to as a difference in gain or simply a gain difference, in the decibel domain.
In other words, the sound volume ratio in the present disclosure is essentially the ratio between the amplitudes of signals; thus, the sound volume ratio may be expressed as a loudness ratio, a volume ratio, an amplitude ratio, a sound level ratio, a sound intensity ratio, a gain ratio, or the like. Furthermore, when the unit of sound volume is decibels, it goes without saying that the sound volume ratio in the present disclosure may be rephrased as the sound volume difference.
In the present disclosure, the “sound volume ratio” typically means the gain difference when the sound volume of each of two sounds is expressed in the unit of decibels, and in the examples of the embodiment, the threshold value data is also typically specified by the gain difference expressed in the decibel domain. However, the sound volume ratio is not limited to the gain difference in the decibel domain. When a sound volume ratio that is not expressed by the decibel domain is used, threshold value data specified in the decibel domain may be used by converting the threshold value data into the unit of the sound volume ratio calculated. Alternatively, threshold value data specified beforehand in each unit may be stored in the memory.
In other words, for example, even if a ratio between energy values, power value, or the like is used instead of the sound volume ratio, it is obvious that the algorithm in the present disclosure can be applied to solve the problem of the present disclosure.
The time difference between the arrival of a direct sound and the arrival of a reflected sound is, for example, the time difference between a direct sound arrival time period (arrival time) and a reflected sound arrival time period (arrival time). It should be noted that for simplicity, the time difference between the arrival of a direct sound and the arrival of a reflected sound may be referred to as the time difference between the direct sound and the reflected sound. The time difference between a direct sound and a reflected sound may be the time difference between the times at which each of the direct sound and the reflected sound arrive at the listening position, the difference in the time periods taken until each of the direct sound and the reflected sound arrive at the listening position, or the time difference between the time when emission of the direct sound ends and the time when the reflected sound arrives at the listening position. The methods for calculating these values will be described later.
Selector 1302 selects whether reproducer 1303 is to generate a reflected sound by using information calculated by analyzer 1301 and the threshold value data. To put it differently, selector 1302 determines whether to select a reflected sound as a reflected sound to be generated. To put it still differently, selector 1302 selects which reflected sounds reproducer 1303 is to generate, from a plurality of reflected sounds.
The threshold value data is, for example, a graph having a horizontal axis that indicates the time difference between a direct sound and reflected sounds and a vertical axis that indicates the sound volume ratios of reflected sounds to a direct sound, and is expressed as a boundary (threshold value) that demarcates whether each reflected sound is perceived. For example, the threshold value data may be expressed as an approximation formula that includes the time difference between a direct sound and a reflected sound as a variable, or may be expressed as an arrangement that includes values of time differences between direct sounds and reflected sounds as an index, and corresponding threshold values.
Selector 1302 selects the generation of a reflected sound when, for example, at the time difference between the arrival time of a direct sound and the arrival time of a reflected sound, the sound volume ratio of the arrival time sound volume of the reflected sound to the arrival time sound volume of the direct sound is a value that is larger than a threshold value set with reference to threshold value data. It should be noted that the arrival time sound volume means the sound volume when the sound arrives at the listening position.
To put it differently, the time difference between the arrival time of a direct sound and the arrival time of a reflected sound is the difference in the amount of time taken for the direct sound and the reflected sound to arrive at the listening position. Furthermore, the time difference between the time point at which emission of the direct sound stops and the time point at which the reflected sound arrives at the listening position may be used as the time difference between the direct sound and the reflected sound. In this case, threshold value data that is different from the threshold value data determined by using, as a reference, the time difference between the direct sound arrival time and the reflected sound arrive time may be used, or common threshold value data may be used.
The threshold value data may be obtained from memory 1404 of audio signal processing device 1001, or may be obtained from an external storage device via a communication module. The threshold value data storage method and the threshold value setting method will be described later.
Reproducer 1303 synthesizes the audio signals of direct sounds and the audio signals of reflected sounds selected for generation by selector 1302.
Specifically, reproducer 1303 processes the inputted audio signals to generate direct sounds, based on information on the direct sound arrival time and the direct sound arrival time sound volume calculated by analyzer 1301. Furthermore, reproducer 1303 processes the inputted audio signals to generate reflected sounds, based on information on the reflected sound arrival time and the reflected sound arrival time sound volume pertaining to the reflected sounds selected by selector 1302. Then, reproducer 1303 synthesizes and outputs the direct sounds and reflected sounds that were generated.
[Operation Example of Renderer]
FIG. 8 is a flowchart illustrating an operation example of audio signal processing device 1001. FIG. 8 illustrates the processing performed mainly by renderer 1300 of audio signal processing device 1001.
In the analysis processing of the input signal (S101 in FIG. 8), analyzer 1301 analyzes the input signal inputted into audio signal processing device 1001 to detect direct sounds and reflected sounds that may be generated in the sound space. The reflected sounds detected here are candidates for the reflected sounds to be selected by selector 1302 as the reflected sounds to be ultimately generated by reproducer 1303. Furthermore, analyzer 1301 analyzes the input signal to calculate information necessary for generating direct sound and reflected sound, and information necessary for selecting the reflected sounds to be generated.
First, the characteristics of each of the direct sound and the reflected sound are calculated. Specifically, the arrival time period and the arrival time sound volume when each of the direct sound and the reflected sound arrive at the listener are calculated. When a plurality of objects are present in the sound space as reflection objects, reflected sound characteristics with respect to each of the plurality of objects are calculated.
The direct sound arrival time period (td) is calculated based on the direct sound arrival path (pd). The direct sound arrival path (pd) is a path that connects position information(S) (xs, ys, zs) of a sound source object with position information A1 (xa, ya, za) of the listener. The direct sound arrival time period (td) is a value obtained by dividing the length of the path that connects position information(S) (xs, ys, zs) with position information A1 (xa, ya, za), by the speed of sound (approximately 340 m/sec).
For example, the path length (X) is determined by the expression ((xs−xa){circumflex over ( )}2+(ys−ya){circumflex over ( )}2+(zs−za){circumflex over ( )}2){circumflex over ( )}0.5. The sound volume attenuates in inverse proportion to the distance. Thus, when the sound volume at position information S (xs, ys, zs) of a sound source object is denoted by N and the unit distance is denoted by U, the direct sound arrival time sound volume (Id) is determined by the expression Id=N*U/X.
Sound volume N at the sound source position may be the reference sound volume described above.
The reflected sound arrival time period (tr) is calculated based on the reflected sound arrival path (pr). The reflected sound arrival path (pr) is a path that connects the position of the sound image of a reflected sound with position information A1 (xa, ya, za).
Note that the position of the sound image of the reflected sound may be derived by using, for example, a “mirror image method” or a “ray tracing method”, or by using any other method for deriving sound image positions. The mirror image method is a method that simulates a sound image by assuming that a reflected wave on the wall in a room has a mirror image in a position symmetrical to the sound source relative to the wall, and that sound waves are emitted from the position of that mirror image. The ray tracing method is a method that simulates, for example, an image (sound image) observed at a certain point by tracing waves that are transmitted in a linear manner, such as light rays or sound rays.
FIG. 9 is a diagram illustrating a comparatively distant positional relationship between a listener and an obstacle object. FIG. 10 is a diagram illustrating a comparatively close positional relationship between a listener and an obstacle object. In other words, each of FIG. 9 and FIG. 10 illustrate an example in which the sound image of a reflected sound is formed in a position symmetrical to the sound source position, with a wall interposed therebetween. Based on such a relationship, by determining the position of the sound image of the reflected sound on the x-, y-, and z-axes, the reflected sound arrival time period can be determined in the same manner as the method for calculating the direct sound arrival time period.
The reflected sound arrival time period (tr) is a value obtained by dividing the length of the path that connects the position of the sound image of a reflected sound with position information A1 (xa, ya, za), by the speed of sound (approximately 340 m/sec). The sound volume attenuates in inverse proportion to the distance. Thus, when the sound volume at the sound source position is denoted by N, the unit distance is denoted by U, and the attenuation rate of the sound volume at the reflection is denoted by G, the reflected sound arrival time sound volume (Ir) is determined by the expression Ir=N*G*U/Y.
As described above, attenuation rate G may be expressed as a real number greater than or equal to 0 and less than or equal to 1, or may be expressed as a negative decibel value. In this case, the sound volume of the signal as a whole attenuates by the amount of G. Furthermore, the attenuation rate may be set for each frequency band included in a plurality of frequency bands. In this case, analyzer 1301 multiplies each frequency component of the signal by the specified attenuation rate. Furthermore, in order to reduce the amount of computation, analyzer 1301 may, by using, as an overall attenuation rate, a representative value, an average value, or the like of a plurality of attenuation rates of a plurality of frequency bands, cause the sound volume of the signal as a whole to attenuate by that amount.
Next, analyzer 1301 calculates the sound volume ratio (L), which is the ratio of the reflected sound arrival time sound volume (Ir) to the direct sound arrival time sound volume (Id), and the time difference (T) between the direct sound and the reflected sound, each of the sound volume ratio (L) and the time difference (T) being required for selection of the reflected sound to be generated.
The sound volume ratio (L), which is the ratio of the above-described Ir to the direct sound arrival time sound volume (Id), is, for example, a value obtained by dividing the reflected sound arrival time sound volume (Ir) by the direct sound arrival time sound volume (Id), and is determined by: L=(N*G*U/Y)/(N*U/X)=G*X/Y. Since the value to be determined is a sound volume ratio, the values of N and U may be any predetermined values.
The time difference (T) between a direct sound and a reflected sound may be, for example, the time difference between the time periods each of the direct sound and the reflected sound take to arrive at the listening position. For example, the time difference (T) between the time periods taken for each of a direct sound and a reflected sound to arrive at the listening position is determined by T=tr−td.
Furthermore, the time difference (T) may be the difference between the times at which each of a direct sound and a reflected sound arrive at the listening position. Moreover, the time difference (T) may be the time difference between the time at which the emission of the direct sound ends and the time at which the reflected sound arrives at the listening position. In other words, the time difference (T) may be the time difference, at the listening position, between the time at which the direct sound ends and the time at which the reflected sound begins.
Next, in reflected sound selection processing (S102 in FIG. 8), selector 1302 selects whether reproducer 1303 is to generate a reflected sound calculated by analyzer 1301. To put it differently, selector 1302 determines whether to select a reflected sound as a reflected sound to be generated. When there are a plurality of reflected sounds, selector 1302 selects whether to generate each reflected sound. As the result of selecting whether to generate each reflected sound, selector 1302 may select one or more reflected sounds to be generated from the plurality of reflected sounds, or may select one reflected sound to be generated.
Note that selector 1302 may select reflected sounds to which other processing is to be applied, not limited to generation processing. For example, selector 1302 may select reflected sounds to which binaural processing is to be applied. Furthermore, selector 1302 fundamentally selects only the one or more reflected sounds that are to be processed. However, selector 1302 may select only one or more reflected sounds that are not to be processed. Processing may then be applied to the one or more reflected sounds that were not selected.
For example, the selection of reflected sounds may be performed based on the sound volume ratio (L) and the time difference (T) calculated by analyzer 1301. Due to the selection processing being performed based on the time difference (T) between direct sounds and reflected sounds, it is possible to more appropriately select reflected sounds that have a large degree of influence on the listener's perception, in comparison to when performing the selection processing based only on the sound volume difference between direct sounds and reflected sounds.
Specifically, the selection of whether to generate a reflected sound is performed by comparing, to a preset threshold value, the sound volume ratio of a reflected sound to a direct sound, the sound volume ratio corresponding to the time difference between the direct sound and the reflected sound. The threshold value is set with reference to the threshold value data. The threshold value data is an indicator indicating the boundary that demarcates whether a reflected sound corresponding to a direct sound is perceived by the listener, and is defined as the ratio between the direct sound arrival time sound volume (Id) and the reflected sound arrival time sound volume (Ir).
Note that the threshold value corresponds to a value expressed by, e.g., a numerical value determined based at the time difference (T). The threshold value data corresponds to the relationship between the time difference (T) and a threshold value, and corresponds to table data or a relational expression used for identifying or calculating the threshold value at the time difference (T). The format and type of the threshold value data is not limited to table data or a relational expression.
FIG. 11 is a diagram illustrating relationships between time differences between direct sounds and reflected sounds, and threshold values. For example, threshold value data of predetermined sound volume ratios may be referenced for each value of the time difference between a direct sound and a reflected sound, as illustrated in FIG. 11. Alternatively, threshold value data obtained by, e.g., interpolating or extrapolating from the threshold value data illustrated in FIG. 11 may be referenced.
Furthermore, the threshold value of the sound volume ratio at the time difference (T) calculated by analyzer 1301 is identified from the threshold value data. Moreover, selector 1302 determines whether to select a reflected sound as a reflected sound to be generated based on whether the sound volume ratio (L) of the reflected sound to the direct sound calculated by analyzer 1301 exceeds the threshold value.
Due to performing the selection processing by using the threshold value data of the sound volume ratio that is predetermined for each value of the time difference between a direct sound and a reflected sound, selection processing that considers post-masking or the precedence effect can be achieved. The type, format, storage method, setting method, and the like of the threshold value data will be described in detail later.
Next, in the generation processing of direct sounds and reflected sounds (S103 in FIG. 8), reproducer 1303 generates and synthesizes the audio signals for direct sounds and the audio signals for reflected sounds that have been selected by selector 1302 as reflected sounds to be generated.
The audio signals for direct sounds are generated by applying the direct sound arrival time period (td) and the direct sound arrival time sound volume (Id) calculated by analyzer 1301 to the sound data for the sound source objects included in the input signal. Specifically, processing in which the sound data is delayed by the amount of the direct sound arrival time period (td) and multiplied by the direct sound arrival time sound volume (Id) is performed. The processing to delay the sound data is processing in which the position of the sound data is moved forward or backward on the time axis. Processing in which the sound data is delayed may be applied without causing the sound quality to deteriorate, such as was disclosed in PTL 2.
The audio signals for reflected sounds are, similarly to the direct sounds, generated by applying the reflected sound arrival time period (tr) and the reflected sound arrival time sound volume (Ir) calculated by analyzer 1301 to the sound data for the sound source objects.
However, the reflected sound arrival time sound volume (Ir) in the generation of reflected sounds differs from the arrival time sound volume of the direct sounds in that the arrival time sound volume of the reflected sounds is a value to which attenuation rate G of the sound volume in the reflection has been applied. G may be an attenuation rate that is applied globally to all frequency bands. Alternatively, in order to reflect the biases of frequency components generated by reflection, the reflectance may be defined for each predetermined frequency band. In this case, the processing to apply the reflected sound arrival time sound volume (Ir) may be performed as frequency equalizer processing, which is processing that involves multiplying each band by the attenuation rate.
In the above example, for each of the direct sounds and the reflected sound candidates, the path length when arriving at the listener is calculated. Furthermore, the arrival time period and the arrival time sound volume are calculated based on each path length. The selection processing of the reflected sound candidates is then performed based on the time differences and the sound volume ratios of these.
Note that as a different example, the selection processing may be performed based on the path lengths when each of the direct sound and the reflected sound arrive at the listener, and the calculation of the arrival time period and the arrival time sound volume of each of the direct sound and the reflected sound and the calculation of the time difference and the sound volume ratio may be omitted. In this case, threshold values according to path length differences may be determined beforehand with respect to path length ratios. Then, selection processing may be performed based on whether the path length ratio calculated is greater than or equal to the threshold value according to the path length difference calculated. This makes it possible to perform selection processing based on path length differences that correspond to time differences, while reducing the amount of computation.
Furthermore, a parameter that indicates sound propagation speed or a parameter that has an impact on the sound propagation speed parameter may be used in addition to the path length difference.
(Details of Selection Processing)
The selection processing that determines whether reflected sounds are generated will be explained in detail.
The selection of a reflected sound is performed by comparing, with the sound volume ratio (L) calculated by analyzer 1301, the threshold value determined for the sound volume ratio, which is the ratio of the reflected sound arrival time sound volume to the direct sound arrival time sound volume, at the time difference (T) between the direct sound and the reflected sound. For example, of threshold values of sound volume ratios that were determined beforehand for each value of a time difference between a direct sound and a reflected sound, the threshold value of the sound volume ratio at the time difference (T) between the direct sound and the reflected sound calculated by analyzer 1301 is referenced. Then, determination of whether the reflected sound is selected as a reflected sound to be generated is made based on whether the sound volume ratio (L) calculated by analyzer 1301 exceeds the threshold value.
The time difference (T) may be any of, for example, the difference in the times at which each of a direct sound and a reflected sound arrive at the listening position, the time difference between the time periods taken when each of a direct sound and a reflected sound arrive at the listening position, or the time difference between the time point when emission of a direct sound stops and the time point when a reflected sound arrives at the listening position. Here, the direct sound end time may be determined by adding the duration of a direct sound to the arrival time of the direct sound.
The threshold value data may be determined based on the minimum time difference at which the perception of the listener is able to detect the divergence of two sounds due to an action of the auditory nerve or a cognitive effect in the brain, and more specifically due to the precedence effect, described later, the temporal masking phenomenon, described later, or a combination of both. Specific numerical values may be derived from research results into the temporal masking effect, the precedence effect, the echo detection limit, etc. that are already known, or may be determined by an auditory test performed with the premise of application in the virtual space.
FIG. 12A, FIG. 12B, and FIG. 12C are diagrams illustrating examples of threshold value data setting methods. As illustrated in FIG. 12A, FIG. 12B, and FIG. 12C, the threshold value data represents the boundaries (threshold values) determining whether reflected sound is perceived or not perceived, in a graph having a horizontal axis that indicates the time difference between direct sound and reflected sound and a vertical axis that indicates the sound volume ratio of the reflected sound to the direct sound.
The threshold value data may be expressed by an approximation formula that includes the time difference between direct sound and reflected sound as a variable. Furthermore, as illustrated in FIG. 11, the threshold value data may be stored in the domain of memory 1404 as an arrangement of an index of time differences between direct sounds and reflected sounds, and threshold values corresponding to the index.
Note that when the height of a line parallel to the horizontal axis in Example 4 in FIG. 12C (the minimum audible limit) is used as the threshold value, what is compared to the threshold value is not the sound volume ratio (L) between a direct sound and a reflected sound, but the sound volume of the reflected sound itself. The reason for this is that the threshold value indicates the sound volume of the boundary demarcating whether a sound is perceivable to the listener, and is a threshold value for determining a sound having a lower sound volume than the threshold value to be a sound that is not to be reproduced. In other words, the threshold value corresponding to the minimum audible limit is not a threshold value with respect to the ratio of the sound volume of a reflected sound to the sound volume of a direct sound.
When the minimum audible limit is used as the threshold value, the threshold value is constant regardless of the time difference (T); thus, the time difference (T) need not be calculated.
Note that when a plurality of reflected sounds are generated in the analysis processing (S101 in FIG. 8), the selection processing may be performed on all of the reflected sounds, or the selection processing may be performed on only the reflected sounds having high evaluation values based on the evaluation values derived for each reflected sound by means of a preset evaluation method. Here, the evaluation value of a reflected sound corresponds to the sensory level of importance of the reflected sound. Note that the evaluation value being high corresponds to the evaluation value being large, and these expressions may be used interchangeably.
Selector 1302 may calculate an evaluation value for each reflected sound by an evaluation method set beforehand based on, for example, the sound volume of the sound source, the visual properties of the sound source, the positionality of the sound source, the visual properties of the reflection object (the obstacle object), the geometrical relationship between the direct sound and the reflected sound, and/or the like.
Specifically, the evaluation value may become higher as the sound volume of the sound source is greater. Furthermore, in order to cause visual positioning and acoustic positioning to match each other, the evaluation value may be high when a sound source object or a reflection object (obstacle object) is visible from the listener, or when the positionality of a sound source object is high.
Moreover, the size of the arrival angle formed by a direct sound and a reflected sound and the difference between the arrival time periods of a direct sound and a reflected sound greatly affect the grasping of the space. Thus, the evaluation value may be high when the size of the angle formed by the arrival of a direct sound and the arrival of a reflected sound is large, or when the difference between the arrival time periods of a direct sound and a reflected sound is large.
The sound volume information on the sound source may indicate the reference sound volume determined for each content, a temporal transition of sound volume, or both.
For example, when the virtual space is a virtual conference room and the direct sound is a speaking voice, the sound volume transitions intermittently over short periods of time. In other words, sound portions and silent portions occur alternately. Furthermore, when the virtual space is a concert hall and the direct sound is the performance of a musical piece, the sound volume is maintained over a certain duration of time. Moreover, when the virtual space is a battlefield and the direct sound is an explosion sound, the sound volume becomes large for only an instant and then continues to be silent or in a quiet state thereafter.
In this way, the sound volume information on the sound source may include not only information on the reference sound volume corresponding to the volume setting when the sound is radiated into the virtual space, but also information on the transition of the sound magnitude.
The information on the transition may be expressed by data indicating frequency characteristics in chronological order. The information on the transition may be expressed by data indicating the duration of a sound interval. The information on the transition may be expressed by data indicating the chronological order of durations of sound intervals and durations of silent intervals. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration for which the amplitude of the sound signal can be considered stationary (can be considered approximately constant) and the amplitude value of said signal during that duration.
The information on the transition may be expressed by data of a duration during which the frequency characteristics of the sound signal can be considered stationary. The information on the transition may be expressed by, for example, data that enumerates, in chronological order, a plurality of sets of a duration for which the frequency characteristics of the sound signal can be considered stationary and the frequency characteristics during that duration.
Furthermore, efforts to use temporal transitions in the frequency characteristics of signals for the acoustic processing of virtual spaces have been conventionally widely performed (PTL 1, etc.). When considering such conventional techniques, it goes without saying that the above-described sets may be sets that include durations of time in which the frequency characteristics are constant and those frequency characteristics.
The geometrical relationships may be positional relationships between a sound source, the listener, and a reflection object in a virtual space. The path lengths of each of the direct sound and the reflected sound arriving can be geometrically calculated based on these relationships. Therefore, using the relationship that sound volume is inversely proportional to distance, it is possible to calculate the reference sound volume of a reflected sound with respect to the reference sound volume of a direct sound.
A reflection coefficient of the reflection object may be used for calculation of the reference sound volume of a reflected sound. Furthermore, a generally used typical value may be used as the reflection coefficient. On the other hand, when special conditions are present, such as the reflection object being covered by a sound-absorbing material or the like, a specially assigned reflection coefficient may be used as the reflection coefficient of the reflection object.
Each reflected sound may be evaluated based on the sound volume of the reflected sound. The sound volume of the reflected sound may, as described above, be determined from the geometrical relationship between the direct sound and the reflected sound, and an indicator assigned to the reflection object. The reflected sound may be evaluated by comparing that sound volume with a predetermined threshold value.
Furthermore, information that indicates a temporal transition in the sound volume of the sound source may be reflected in the evaluation. For example, in a case in which the information that indicates a temporal transition in the sound volume of the sound source indicates the duration of a sound interval, the reflected sound evaluation value may be maintained as-is when the time is within the sound interval. On the other hand, when the time is outside of the sound interval, processing to reduce or make zero the evaluation value of the reflected sound may be performed, even if the reference sound volume of the reflected sound exceeds the threshold value.
Alternatively, the information that indicates a temporal transition in the sound volume of the sound source may be data that lists, in chronological order, a plurality of sets of a duration during which it is considered that the amplitude of the sound signal is largely constant and amplitude values of the signal during that period. In that case, processing may be performed such that reflected sounds are evaluated by changing the reference sound volume of the reflected sounds in coordination with changes in the amplitude values in the data.
Furthermore, both the information on the reference sound volume and the information on the temporally transitioning sound volume may be used as the information indicating the sound volume of a direct sound. For example, an evaluation value can be calculated based on the information on the reference sound volume, and then the evaluation value can be corrected by using the information on the transitioning sound volume.
In the evaluation of reflected sounds, all of the methods described above may be performed, or only some of the methods may be performed. For example, reflected sounds may be evaluated using a plurality of evaluation methods, or reflected sounds may be evaluated using one evaluation method.
When reflected sounds are evaluated using a plurality of evaluation methods, whether to select each reflected sound may be determined based on an evaluation value comprehensively determined by the plurality of evaluation methods, or may be determined based on the evaluation value of each of the plurality of evaluation methods.
When whether to select each reflected sound is determined based on each of the plurality of evaluation methods, audio signal processing device 1001 may select a sound when all of the plurality of evaluation results based on the plurality of evaluation methods indicate selecting the sound. Alternatively, audio signal processing device 1001 may select a sound when any one of the plurality of the evaluation results based on the plurality of evaluation methods indicates selecting the sound.
Furthermore, for example, an order of priority may be assigned to the first to third evaluation methods. When it is determined by the first evaluation method that a sound is not to be selected, audio signal processing device 1001 may then make a final determination that the sound is not to be selected, without depending on the determination results in the second and third evaluation methods. Moreover, when, although it has been determined by one of the second and third evaluation methods that a sound is not to be selected, it has been determined by the other of the second and third evaluation methods that the sound is to be selected, audio signal processing device 1001 may make a final determination that the sound is to be selected.
Furthermore, the selection processing and the evaluation processing may be performed independently of each other, or only one of these may be performed. Moreover, the evaluation processing may be performed only on reflected sounds that have been determined to be selected in the selection processing, and whether to select each reflected sound may be redetermined in the evaluation processing. Alternatively, the evaluation processing may be performed only on reflected sounds that have been determined to not be selected in the selection processing, and whether to select each reflected sound may be redetermined in the evaluation processing.
The selection processing described above can be interpreted as processing in which a reflected sound is selected in accordance with a characteristic of a direct sound. For example, in processing in which a reflected sound is selected in accordance with a characteristic of a direct sound, the threshold value used in selection of the reflected sound is set or adjusted in accordance with a characteristic of the direct sound. Alternatively, the evaluation value used in the selection of a reflected sound may be calculated based on one or more of, for example, the sound volume of the sound source, the visual properties of the sound source, the positionality of the sound source, the visual properties of the reflection object (the obstacle object), the geometrical relationship between the direct sound and the reflected sound, and/or the like.
Furthermore, the processing in which a reflected sound is selected based on a characteristic of a direct sound is not limited to processing in which the threshold value is set or adjusted in accordance with a characteristic of the direct sound and processing in which the evaluation value used for selection of the reflected sound to be processed is calculated, and other processes may be performed. Furthermore, even when performing the processing in which the threshold value is set or adjusted in accordance with a characteristic of the direct sound or the processing in which the evaluation value used in selection of the reflected sounds to be processed is calculated, the processing may be partially changed, or new processing may be added.
Note that setting the threshold value may include adjusting the threshold value, changing the threshold value, and the like.
[Threshold Value Setting Method]
The threshold value data used in the selection processing may be set with reference to the value of an echo detection limit based on a known precedence effect or a masking threshold value based on the post-masking effect.
The precedence effect is a phenomenon in which, when sounds are heard from two locations, it is perceived that the sound source is present at the location from which the first sound was heard. If two short sounds fuse together to be heard as one sound, the position (localization position) from which the overall sound is heard is, for the most part, determined by the position of the first sound. The echo detection limit is a phenomenon that occurs due to the precedence effect, and is the minimum time difference at which the listener's perception detects the divergence of two sounds.
In Example 2 of FIG. 12C, the horizontal axis corresponds to the arrival time period of reflected sound (echo), and specifically corresponds to the delay time period from the arrival time of direct sound to the arrival time of reflected sound. The vertical axis corresponds to the sound volume ratio of detectable reflected sound to direct sound, and specifically corresponds to the threshold value that determines whether reflected sound that has arrived with a delay time period is detectable.
FIG. 13 is a diagram illustrating an example of a threshold value setting method. The horizontal axis in FIG. 13 corresponds to the arrival time period of reflected sound, and specifically corresponds to the time differences (T) between direct sound and reflected sound. The vertical axis in FIG. 13 corresponds to the sound volume of reflected sound. Specifically, the vertical axis in FIG. 13 may correspond to the sound volume (sound volume ratios) of reflected sound determined in relation to direct sound, or may correspond to the sound volume of reflected sound determined absolutely without depending on the sound volume of the direct sound.
For example, when, as illustrated in FIG. 9, the listener and an obstacle object are comparatively far from each other, the arrival time period of the reflected sound becomes longer, and, as illustrated in C in FIG. 13, the threshold value is set to be low. As a result, in the case of FIG. 9, the reflected sound is generated. On the other hand, when, as illustrated in FIG. 10, the listener and the obstacle object are comparatively close to each other, the arrival time period of the reflected sound is shorter than that in the case of FIG. 9, and as illustrated in B in FIG. 13, the threshold value is set to be high. As a result, in the case of FIG. 10, the reflected sound is not generated.
Furthermore, the threshold value data may be stored in memory 1404, obtained from memory 1404 at the time of the selection processing, and used in the selection processing.
FIG. 14 is a flowchart illustrating an example of selection processing. First, selector 1302 specifies a reflected sound detected by analyzer 1301 (S201). Selector 1302 then detects the sound volume ratio (L) of the reflected sound to the direct sound, and the time difference (T) between the direct sound and the reflected sound (S202 and S203).
The time difference (T) may be any of, for example, the time difference between the time periods each of the direct sound and the reflected sound take to arrive at the listening position, the time difference between the direct sound arrival time and the reflected sound arrival time, and the time difference between the time when emission of the direct sound ends and the time when the reflected sound arrives at the listening position. Here, an example will be described based on the time difference between the direct sound arrival time and the reflected sound arrival time.
Specifically, based on: the position information on the sound source object and the listener; and the position information and geometry information on the obstacle object, selector 1302 calculates the difference between the length of the path of the direct sound and the length of the path of the reflected sound. By dividing the difference between the lengths by the speed of sound, selector 1302 then detects the time difference (T) between the time when the direct sound arrives at the listener's position and the time when the reflected sound arrives at the listener's position.
The sound volume when arriving at the listener attenuates, with respect to the sound volume of the sound source, in proportion to the distance to the listener (in inverse proportion to the distance). Therefore, the sound volume of the direct sound is obtained by dividing the sound volume of the sound source by the length of the path of the direct sound. The sound volume of the reflected sound is obtained by dividing the sound volume of the sound source by the length of the path of the reflected sound, and then further multiplying by the attenuation rate assigned to the virtual obstacle object. Selector 1302 detects the sound volume ratio by calculating the ratio between these sound volumes.
Furthermore, using the threshold value data, selector 1302 identifies the threshold value corresponding to the time difference (T) (S204). Selector 1302 then determines whether the sound volume ratio (L) detected is greater than or equal to the threshold value (S205).
When the sound volume ratio (L) is greater than or equal to the threshold value (“Yes” in S205), selector 1302 selects the reflected sound as a reflected sound to be generated (S206). When the sound volume ratio (L) is less than the threshold value (“No” in S205), selector 1302 skips selecting the reflected sound as a reflected sound to be generated (S207). That is, in this case, selector 1302 determines the reflected sound to be a reflected sound that is not to be generated, i.e., a reflected sound to be culled.
Subsequently, selector 1302 determines whether there are any unspecified reflected sounds (S208). If there are unspecified reflected sounds (“Yes” in S208), selector 1302 repeats the above-described processing (S201 to S207). If there are no unspecified reflected sounds (“No” in S208), selector 1302 ends the processing.
This selection processing may be performed on all of the reflected sounds generated in the analysis processing, or may be performed on only the reflected sounds for which the above-described evaluation value is high.
[Details of Threshold Value Storing Method]
The threshold value data according to the present embodiment is stored in memory 1404 of audio signal processing device 1001. The format and type of the threshold value data to be stored may be any format and any type. When threshold values having a plurality of formats and a plurality of types are stored, in the selection processing, the format and the type of the threshold values to be used in the selection processing of the reflected sounds may be decided. The method for determining which items of threshold value data to use in the selection processing will be described later.
Furthermore, a plurality of formats and a plurality of types of threshold value data may be stored in combination. The combined threshold value data may be read from spatial information managers 1201 and 1211 to set the threshold values to be used in the selection processing. Note that the threshold value data to be stored in memory 1404 may be stored in spatial information managers 1201 and 1211.
For example, the threshold value data may be stored as threshold values at each time difference, so as to plot a line between the threshold values as illustrated in [Example 1] and [Example 2] of FIG. 12C.
Furthermore, the threshold value data may be stored as table data in which, as illustrated in FIG. 11, the threshold values and the time differences (T) are associated with each other. In other words, the threshold value data may be stored as table data that includes the time differences (T) as an index. Naturally, the threshold values illustrated in FIG. 11 are examples, and the threshold values are not limited to the examples in FIG. 11. Furthermore, the threshold values may be approximated by functions that include the time differences (T) as variables, and coefficients of the functions may be stored, without storing the threshold values themselves. Moreover, a plurality of approximation expressions may be combined and stored.
For example, the threshold value data may be expressed by a formula such as that shown below, where timeDiff denotes the time difference (T), and gainThresh denotes the threshold value.
The threshold value is defined only over the time range in which the precedence effect is considered to occur. When the time difference is outside of that time range (in the above formula, a value of 1 ms or less or 40 ms or greater), the determination may be performed not by gainThresh, but solely by a threshold value representing the minimum sound volume for reproduction in the virtual space, described later.
Experiments by the present inventors have clarified that in the time range in which the precedence effect is considered to occur, the threshold value may be approximated by a function that is convex upward. The above-described formula is an example of an approximation formula generated based on these experiments.
Information on a relational expression that indicates the relationship between time differences (T) and threshold values may be stored in memory 1404. In other words, an expression that includes the time difference (T) as a variable may be stored. The threshold values of the time differences (T) may be approximated by a straight line or a curved line, and a parameter that indicates the geometrical shape of the straight line or the curved line may be stored. For example, when the geometrical shape is a straight line, the start point and the slope for expressing the straight line may be stored.
Furthermore, the threshold value data may be stored having the type and format thereof defined for each characteristic of direct sound. Moreover, parameters for adjusting threshold values based on a characteristic of the direct sound and using the threshold values in the selection processing may be stored. Processing to adjust threshold values in accordance with a characteristic of the direct sound and use the threshold values in the selection processing is described later, as a variation of the threshold value setting method.
As an example in which a plurality of types of threshold value data are stored in combination, as illustrated in [Example 3] in FIG. 12C, for each time difference (T), the larger value of the masking threshold value and the echo detection limit threshold value may be stored. As illustrated in [Example 4] in FIG. 12C, for each time difference (T), the larger value of the minimum sound volume for reproduction in a virtual space and the echo detection limit threshold value may be stored.
The combination of the plurality of types of the threshold value data is not limited to these. For example, in a plurality of items of threshold value data, information on the maximum value may be stored for each time difference (T).
Furthermore, in the above description, the information on threshold values has time period items as a one-dimensional index. The information on threshold values may have a two-dimensional or three-dimensional index that further includes variables related to the direction of arrival.
FIG. 15 is a diagram illustrating relationships between directions of direct sounds, directions of reflected sounds, time differences, and threshold values. For example, as illustrated in FIG. 15, threshold values pre-calculated in accordance with the relationship between the direct sound direction (θ), the reflected sound direction (γ), the time difference (T), and the sound volume ratio (L) may be stored.
The direct sound direction (θ) corresponds to the angle, with respect to the listener, of the direction of arrival of a direct sound. The reflected sound direction (γ) corresponds to the angle, with respect to the listener, of the direction of arrival of a reflected sound. Here, the direction in which the listener is facing is defined as 0 degrees. The time difference (T) corresponds to the difference, until reaching the listening position, of the direct sound arrival time period and the reflected sound arrival time period. The sound volume ratio (L) corresponds to the sound volume ratio of the reflected sound arrival time sound volume to the direct sound arrival time sound volume.
Naturally, the threshold values illustrated in FIG. 15 are examples, and the threshold values are not limited to the examples in FIG. 15. Furthermore, in FIG. 15, mainly threshold values when the angle (θ) of the direct sound arrival direction is 0 degrees are exemplified. However, threshold values when the direct sound arrival direction (θ) is not 0 degrees are also stored in memory 1404.
Further, in the above description, the threshold values are stored in an arrangement that has the angle (θ) of the direct sound (more specifically, the angle (θ) of the direct sound arrival direction) and the angle (γ) of the reflected sound (more specifically, the angle (γ) of the reflected sound arrival direction) as independent variables or indexes. However, the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction need not be used as independent variables.
For example, the angular difference between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be used. This angular difference corresponds to the angle formed between the direct sound arrival direction and the reflected sound arrival direction, and may be expressed as the arrival angle between a direct sound and a reflected sound.
FIG. 16 is a diagram illustrating relationships between angular differences, time differences, and threshold values. For example, threshold values pre-calculated by using, as a variable, the angular difference (φ) between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be stored as in the example illustrated in FIG. 16. Naturally, the threshold values illustrated in FIG. 16 are examples, and the threshold values are not limited to the examples in FIG. 16.
In the example in FIG. 16, the number of variables used for deriving threshold values may be reduced. Thus, it is possible to reduce the number of threshold values stored in memory 1404. Therefore, it is possible to decrease the amount of data stored in memory 1404.
Furthermore, when the angular difference (φ) between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction is used, the threshold value data may be stored in a two-dimensional arrangement. Moreover, in the selection processing, the difference between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be calculated by using a three-dimensional arrangement.
The method for selecting reflected sounds using threshold values based on the directions of arrival will be described later.
[First Variation of Threshold Value Setting Method]
In the examples in FIG. 12A, FIG. 12B, and FIG. 12C, threshold values in a plurality of formats and of a plurality of types may be stored in spatial information managers 1201 and 1211. Then, of the threshold values having a plurality of formats and a plurality of types, the format and the type of the threshold values to be used in the selection processing of the reflected sounds may be decided. Specifically, as illustrated in Example 3 of FIG. 12C, in the time differences (T) corresponding to the reflected sound arrival times, the largest threshold value may be adopted.
Moreover, as illustrated in Example 4, the masking threshold value, the echo detection limit threshold value, and a threshold value indicating the minimum sound volume for reproduction in the virtual space may be stored. Then, the largest threshold value at the time difference (T) corresponding to the reflected sound arrival time may be adopted.
[Second Variation of Threshold Value Setting Method]
As another example of the threshold value setting method, a method for setting threshold values in accordance with a characteristic of direct sounds will be described.
FIG. 17 is a block diagram illustrating another configuration example of renderer 1300 illustrated in FIG. 7. Renderer 1300 in FIG. 17 is different from renderer 1300 in FIG. 7 in the respect that renderer 1300 in FIG. 17 includes threshold value adjuster 1304. The description other than threshold value adjuster 1304 is the same as the matters described regarding FIG. 7, and has thus been omitted.
Threshold value adjuster 1304 selects, from the threshold value data, threshold values that are to be used by selector 1302, based on information indicating a characteristic of an audio signal. Alternatively, threshold value adjuster 1304 may adjust the threshold values included in the threshold value data, based on the information indicating a characteristic of the audio signal.
The information indicating a characteristic of the audio signal may be included in the input signal. Then, threshold value adjuster 1304 may obtain the information indicating a characteristic of the audio signal from the input signal. Alternatively, analyzer 1301 may derive a characteristic of the audio signal by analyzing the audio signal included in the input signal accepted by analyzer 1301, and output the information indicating the characteristic of the audio signal to threshold value adjuster 1304.
The information indicating a characteristic of the audio signal may be obtained before starting the rendering processing, or may be obtained each time during rendering.
Furthermore, threshold value adjuster 1304 need not be included in audio signal processing device 1001; another transmission device may have the role of threshold value adjuster 1304. In this case, analyzer 1301 or selector 1302 may obtain, from the other transmission device via communication I/F 1403, the information indicating a characteristic of the audio signal, the threshold value data corresponding to the characteristic, or information for adjusting the threshold value data in accordance with the characteristic.
FIG. 18 is a flowchart illustrating another example of selection processing. FIG. 19 is a flowchart illustrating yet another example of selection processing. In FIG. 18 and FIG. 19, the threshold value is set in accordance with a characteristic of the direct sound. Specifically, in FIG. 18, threshold value adjuster 1304 identifies a threshold value from the threshold value data, based on the time difference (T) and a characteristic of the audio signal. In FIG. 19, threshold value adjuster 1304 adjusts, based on a characteristic of the audio signal, the threshold value identified from the threshold value data based on the time difference (T).
Hereinafter, the operations of each example will be described. Note that description has been omitted for processes that are shared with the example in FIG. 14.
First, an example of the processing illustrated in FIG. 18 will be described. Here, the threshold value data is stored beforehand in memory 1404 for each nature of direct sound. Accordingly, a plurality of items of threshold value data corresponding to a plurality of natures are stored beforehand in memory 1404. Then, threshold value adjuster 1304 identifies, from the plurality of items of threshold value data, the threshold value data to be used in the selection processing of reflected sounds.
For example, threshold value adjuster 1304 obtains a characteristic of a direct sound based on the input signal (S211). Threshold value adjuster 1304 may obtain a characteristic of the direct sound that is associated with the input signal. Threshold value adjuster 1304 may then identify the threshold value corresponding to the time difference (T) and the characteristic of the direct sound (S212).
Furthermore, as illustrated in FIG. 19, threshold value adjuster 1304 may adjust the threshold value identified by selector 1302, based on the characteristic of the direct sound (S221).
In any of these cases, the information indicating a characteristic of the audio signal, the information for adjusting the threshold value in accordance with the characteristic of the audio signal, or both of these may be included in the input signal. Threshold value adjuster 1304 may adjust the threshold value using one or both of these.
Furthermore, the information indicating a characteristic of the audio signal, the information for adjusting the threshold value, or both of these may be transmitted by another input signal aside from the input signal that includes the audio signal. In this case, information for associating the other input signal aside from the input signal may be included in the input signal that includes the audio signal, or information for associating the other input signal with the input signal may be stored in memory 1404 together with the information on threshold values.
In the examples in FIG. 18 and FIG. 19, the threshold value used in selecting each reflected sound is set in accordance with a characteristic of the direct sound, that is, a characteristic of the audio signal. Threshold value data preset for each characteristic may be used, as in FIG. 18, or the threshold value may be adjusted in accordance with the characteristics of the audio signal, as in FIG. 19. Furthermore, threshold value data parameters may be adjusted in accordance with the characteristic of the audio signal.
Moreover, the operations performed by threshold value adjuster 1304 may be performed by analyzer 1301 or selector 1302. For example, analyzer 1301 may obtain a characteristic of the audio signal. Furthermore, selector 1302 may set threshold values in accordance with the characteristic of the audio signal.
Next, the relationship between the characteristic of the audio signal and the threshold value will be described.
Two short sounds that arrive at the listener's ears in succession are heard as one sound if the time period interval between the two short sounds is sufficiently short. This phenomenon is referred to as the precedence effect. The precedence effect is known to only occur with respect to unconnected sounds, that is, transient sounds (NPL 1). Thus, when an audio signal indicates a stationary sound, the echo detection limit may be set lower than when the audio signal indicates a non-stationary sound.
In other words, the threshold value is set low in accordance with the characteristics of this precedence effect when, for example, a direct sound is a stationary sound. Furthermore, the threshold value may be set lower as the stationarity is greater.
An example of processing when the characteristic of the audio signal is stationary will be explained. First, threshold value adjuster 1304 or analyzer 1301 makes a determination on the stationarity based on the amount of variation in a frequency component of an audio signal accompanying the passage of time. For example, when the amount of variation is small, it is determined that the stationarity is high. Conversely, when the amount of variation is great, it is determined that the stationarity is low. As a result of the determination, a graph indicating the level of stationarity may be set, or a parameter indicating the stationarity in accordance with the amount of variation may be set.
Next, threshold value adjuster 1304 adjusts the threshold value data or the threshold values based on information indicating the stationarity, such as the graph or the parameter indicating the stationarity of the audio signal, and sets the adjusted threshold value data or threshold values as threshold value data or threshold values to be used by selector 1302.
Alternatively, a parameter for setting the threshold value data in accordance with the information indicating direct sound stationarity may be stored beforehand in memory 1404. In this case, threshold value adjuster 1304 may make a determination on the stationarity of the audio signal and set the threshold value data to be used in the selection of reflected sounds, based on the information indicating stationarity and the parameter.
Alternatively, a plurality of parameters for threshold value data may be stored beforehand in memory 1404, corresponding to a plurality of patterns of direct sound stationarity. In this case, threshold value adjuster 1304 may make a determination on the stationarity of the audio signal, select the threshold value data parameter based on the pattern of direct sound stationarity, and set the threshold value data to be used in the selection of reflected sounds, based on the threshold value data parameter.
Note that a determination on the stationarity of an audio signal may be made based on the amount of variation of the frequency component of the audio signal, each time an audio signal is inputted.
Alternatively, a determination on the stationarity of an audio signal may be made based on information indicating stationarity that is pre-associated with the audio signal. In other words, the information indicating audio signal stationarity may be associated with the audio signal and pre-stored in memory 1404. Analyzer 1301 may, each time an audio signal is inputted, obtain information indicating stationarity that is associated with the audio signal. Threshold value adjuster 1304 may then adjust the threshold values based on the information indicating stationarity that is associated with the audio signal.
As another example of threshold values being set in accordance with a characteristic of the audio signal, when an audio signal indicates short sounds (clicking sounds, etc.), the application scope of the echo detection limit may be set shorter than when an audio signal indicates long sounds. This processing is based on the characteristics of the precedence effect.
It is known that due to the precedence effect, two short sounds that arrive at the listener's ears in succession are heard as one sound if the time period interval between the two short sounds is sufficiently short. The upper limit of this time period interval is dependent on the length of the sounds. For example, the upper limit of this time period interval is about 5 ms for clicking sounds, but for complex sounds such as a human voice or music, the upper limit may be 40 ms (NPL 1).
In accordance with the characteristics of this precedence effect, for example, in the case of a sound for which the duration of a direct sound is short, threshold values for short time period lengths are set. Furthermore, threshold values for shorter time period lengths are set as the duration of the direct sound is shorter.
Threshold values for short time period lengths being set means that within a range in which the time difference (T) between a direct sound and a reflected sound is small, threshold values corresponding to an echo detection limit based on the characteristics of the precedence effect are set. Threshold values corresponding to the echo detection limit based on the characteristics of the precedence effect are not set outside of this range. In other words, outside of this range, threshold values are low. Thus, threshold values for short time period lengths being set for short sounds can correspond to low threshold values being set for short sounds.
As another example of threshold values being set in accordance with a characteristic of a direct sound, when a direct sound is an intermittent sound (such as speech), threshold values may be set lower than when a direct sound is a continuous sound (such as music).
For example, when a direct sound corresponds to speech, sound portions and silent portions repeat, and in the silent portions, only the post-masking effect occurs as the masking effect. On the other hand, when the direct sound is a continuous sound such as musical content, the masking effects that occur include both the post-masking effect and a simultaneous masking effect that results from sound occurring at that time. Consequently, the overall masking effect is greater in the case of music, etc. than in the case of speech, etc.
In accordance with masking effect characteristics such as those described above, threshold values may be set higher in the case of music, etc. than in the case of speech, etc. Conversely, threshold values may be set lower in the case of speech, etc. than in the case of music, etc. That is, threshold values may be set to be low when a direct sound has numerous intermittent portions.
As described above, the information indicating a characteristic of a direct sound may be information indicating the stationarity, intermittency, duration, etc. of the direct sound. Furthermore, the information indicating a characteristic of a direct sound may be any combination of these. Furthermore, the information indicating a characteristic of a direct sound may be information indicating the time variation of one of these, or may be information indicating the time variation of any combination of these. That is, the information indicating a characteristic of a direct sound may be information indicating the time variation of the direct sound.
For example, as indicated in the description of the stationarity determination, the information indicating a characteristic of a direct sound may be chronological data on the frequency characteristics. Here, the frequency characteristics may be expressed in a commonly used form such as a gain value per frequency band, a Fourier series with respect to a time axis signal, a linear predictive coding (LPC) coefficient or cepstral coefficient for determining a frequency envelope, or the like.
Furthermore, the information indicating a characteristic of a direct sound may be, as information indicating the intermittency of direct sound, information (an amplitude envelope outline) enumerating, in chronological order, a plurality of sets of: the duration for which the amplitude of a signal is stationary; and the amplitude value of the signal for that duration. Here, the amplitude value may be expressed as a ratio with respect to the reference sound volume.
Furthermore, the information indicating a characteristic of a direct sound may be information on the frequency characteristics of the direct sound. For example, the information indicating a characteristic of a direct sound may be information indicating the stationarity of the frequency characteristics of the direct sound. Specifically the information indicating a characteristic of a direct sound may be information (a spectrogram outline) enumerating, in chronological order, a plurality of sets of: a duration for which the variation in frequency characteristics is small; and the frequency characteristics of the signal for the duration. Here, the sound volume used as the reference for the frequency characteristics may be the reference sound volume.
For example, the information indicating the time variation of a direct sound may be information indicating a direct sound envelope. The information indicating the time variation of a direct sound may be used when the “minimum audible limit” described in [Example 4] of FIG. 12C is the threshold value. The signal compared to the minimum audible limit is the sound volume of a reflected sound.
The sound volume of a reflected sound is obtained by geometrical calculation from information on the positions of the sound source, the listener, and the reflection object. Specifically, the reference sound volume of the reflected sound is obtained with respect to the reference sound volume of the sound source. By adjusting the reference sound volume of the reflected sound using, as the information on a characteristic of a direct sound, information on the sound magnitude transition of the sound source, the sound volume of the reflected sound at each moment can be accurately determined. This is because the variation in the sound volume of the sound source is reflected in the variation in the sound volume of the reflected sound.
By comparing the sound volume of a reflected sound with the threshold value after adjusting the sound volume of the reflected sound, the reflected sounds that are auditorily necessary can be appropriately selected with more accuracy.
It goes without saying that naturally, the same result can be obtained by, without adjusting the reference sound volume of a reflected sound, adjusting the threshold value based on the inverse of the information on the sound magnitude transition of the sound source, and comparing the adjusted threshold value to the reference sound volume of the reflected sound. That is, the reference sound volume of a reflected sound may be adjusted using the information on the sound magnitude transition of the sound source, or the threshold value may be adjusted using the information on the sound magnitude transition of the sound source. The adjustment of the reference sound volume of a reflected sound and the adjustment of a threshold value correspond to each other.
Depending on the composition of the surface of an object that reflects sound, the reflectance of sound (the attenuation rate accompanying the reflection) is different for each frequency band. Accordingly, as described later, the reflectance (attenuation rate) of sound may be associated with each frequency band, for an object that reflects sound. Whether to select the reflected sound can be more accurately determined based on such reflectance information and spectrogram information. For example, processing such as the following is performed.
Specifically, for example, it is indicated by spectrogram information that a high-frequency component is more dominant than a low-frequency component at a certain time interval. Furthermore, for example, it is indicated by sound reflectance information that the reflectance is very low at the high-frequency component, compared to the low-frequency component.
In this case, there is a possibility that even if the amplitude of the sound source signal on the time axis is high, the sound volume of the reflected sound will be low, the sound volume of the reflected sound being obtained by multiplying, by the attenuation rate for each frequency band indicated by the reflectance information, the frequency component indicated by the spectrogram information. Thus, the reflected sound will not be selected.
As described above, the information indicating a characteristic of a direct sound may be information indicating the time variation of the direct sound. For example, the information indicating a characteristic of a direct sound may indicate a value obtained by analyzing the direct sound at a predetermined time length.
Specifically, the information indicating a characteristic of a direct sound may be information obtained by calculating the average energy or the average amplitude of the direct sound for each predetermined time length. Furthermore, the information indicating a characteristic of a direct sound may be information obtained by calculating the energy or average amplitude of the direct sound for each short time analysis length, and calculating the weighted average of the energy or the average amplitude for each long time analysis length that is longer than the short time analysis length.
More specifically, for example, the information indicating the time variation of a direct sound may be information obtained by calculating the energy or the average amplitude of the direct sound for each predetermined short time length (for example, 5 ms; hereinafter, time length frames are expressed as analysis frames). Furthermore, the information indicating the time variation of a direct sound may be information represented by a weighted average of the energy or average amplitude calculated using the past N−1 analysis frames.
When the energy of the n-th analysis frame is expressed by E(n), information I(n) indicating a characteristic of the direct sound is determined in accordance with the following formula.
Here, parameter a(i) denotes a weighting factor. Typically, a(i) is set such that a(i)≥0 and the sum of a(i) is 1. However, the method for setting a(i) is not limited thereto.
Note that information I(n) indicating a characteristic of direct sound is calculated each time 5 ms of direct sound is imported. That is to say, the time variation of information I(n) indicating a characteristic of a direct sound can be calculated with low delay. Thus, this method is suitably applied to an application requiring real-time performance.
Furthermore, information I(n) indicating a characteristic of a direct sound may be determined in accordance with the following formula.
Here, parameter b(i) denotes a weighting factor. Typically, b(i) is set such that b(i)≥0 and the sum of b(i) is 1. However, the method for setting b(i) is not limited thereto.
In this formula, information I(n) indicating a characteristic of a direct sound is recursively determined. Thus, the average energy for a long time length can be calculated with a small amount of computation.
Formula 1 and Formula 2 above can be considered to be filters in which E(n) is the input signal and I(n) is the output signal. In this case, Formula 1 is a moving average (MA) model filter and Formula 2 is an autoregressive (AR) model filter, and both of these have the characteristics of a low-pass filter. Furthermore, an ARMA model filter, in which both of these are combined, may be used.
Note that the method for deriving the information indicating the time variation of a direct sound is not limited to the above-described formulae or filters, and other well-known methods may be used. As described above, the information indicating the time variation of a direct sound indicates a value obtained by analyzing a direct sound at a predetermined time length. A direct sound may be analyzed from perspectives other than average energy.
Furthermore, as described above, the information indicating a characteristic of a direct sound may be information on the frequency characteristics of the direct sound. The information on the frequency characteristics of a direct sound may be information calculated by using the frequency characteristics of the direct sound. For example, the information on the frequency characteristics of a direct sound may be information obtained as the average energy of a low-frequency component, by averaging the low-frequency component of the direct sound at a predetermined analysis length.
Specifically, the low-frequency component of a direct sound can be determined by applying, to a direct sound contained in an analysis frame length, a filter having low-pass characteristics. From the energy or average amplitude of this low-frequency component, the information indicating a characteristic of a direct sound can be derived, similarly to Formula 1 described above.
When the energy of the low-frequency component of the n-th analysis frame is expressed by EL(n), information I(n) indicating a characteristic of the direct sound can be determined in accordance with the following formula.
Here, parameter c(i) denotes a weighting factor. Typically, c(i) is set such that c(i)≥0 and the sum of c(i) is 1. However, the method for setting c(i) is not limited thereto.
Note that information I(n) indicating a characteristic of direct sound is calculated each time 5 ms of direct sound is imported. That is to say, the time variation of information I(n) indicating a characteristic of a direct sound can be calculated with low delay. Thus, this method is suitably applied to an application requiring real-time performance.
Furthermore, similarly to Formula 2, information I(n) indicating a characteristic of a direct sound can be determined in accordance with the following formula.
Here, parameter d(i) denotes a weighting factor. Typically, d(i) is set such that d(i)≥0 and the sum of d(i) is 1. However, the method for setting d(i) is not limited thereto.
In this formula, information I(n) indicating a characteristic of a direct sound is recursively determined. Thus, the average energy for a long time length can be calculated with a small amount of computation.
Formula 3 and Formula 4 above can be considered to be filters in which E(n) is the input signal and I(n) is the output signal. In this case, Formula 3 is a moving average (MA) model filter and Formula 4 is an autoregressive (AR) model filter, and both of these have the characteristics of a low-pass filter. Furthermore, an ARMA model filter, in which both of these are combined, may be used.
In the above description, a filter having low-pass characteristics was used in the method for determining the low-frequency component of a direct sound, but the method for determining the low-frequency component of a direct sound is not limited thereto. Furthermore, the method for deriving the information indicating the time variation of direct sound is not limited to the above-described formulae or filters, and other well-known methods may be used. Furthermore, the spectrum of a direct sound can be calculated by applying frequency conversion to the direct sound. The energy or average amplitude of the low-frequency component of the spectrum can then be calculated.
Furthermore, in the above description, the MA model or the AR model is used for deriving the information indicating the time variation of a direct sound. The coefficients of these models may be predetermined fixed values, or may be variable values, which are values that temporally vary.
Furthermore, the relationship between the analysis frame length and the interval at which the information update threads are created may be as described below.
For example, when the time length of an analysis frame is TA (msec) and the interval at which information update threads are created is TU (msec), the value of N in (Formula 1) and (Formula 3), described above, in the MA filter may be approximately a value given by TU/TA. Furthermore, b(i) and d(i) (1≤i<N) in (Formula 2) and (Formula 4), described above, in the AR filter may be a value that results in the time constant of the filter being about TU (msec). The reason for this setting is that within the information update interval period, convergence of the filter is expected.
On the other hand, when, in the information indicating the time variation of a direct sound with the above-described setting, the value varies too sharply, I(n) may be precalculated. Furthermore, precalculated I(n) may be applied to the selection processing of reflected sounds. For example, in the processing of the t-th time frame, I(t+tau) may be used. Here, tau is a value defined in accordance with the characteristics of filter convergence. When the convergence is slow, the value of tau is large in comparison to when the convergence is fast.
Furthermore, as the information indicating the characteristics of a direct sound, information on auditory masking (frequency masking) calculated from the direct sound may be used. The auditory masking information indicates a threshold value of an amplitude value at a frequency region masked by a direct sound. Processing in which the amplitude values of reflected sounds in the same frequency range are compared to the threshold value, and a reflected sound having an amplitude value lower than the threshold value is not selected may be performed. The amplitude value of a reflected sound in a frequency range may be obtained by analyzer 1301 as information indicating the characteristics of a reflected sound.
When threshold values to be used in selecting reflected sounds are thus set in accordance with the characteristics of direct sound, it is possible to appropriately select reflected sounds that are auditorily necessary, and the characteristics of auditory sensitivity can be effectively reflected in three-dimensional sound reproduction system 1000. Processing to detect characteristics of direct sound, processing to determine threshold values in accordance with the characteristics, and processing to adjust the threshold values in accordance with the characteristics may be performed during the rendering processing, or may be performed before starting the rendering processing.
For example, these processes may be performed, for example, during virtual space creation (during software creation), when starting processing of the virtual space (when launching the software or starting rendering), or when there is an occurrence of an information update thread that periodically occurs in processing of the virtual space. Furthermore, the time of virtual space creation may be when the virtual space is built before starting acoustic processing, may be when information (spatial information) on the virtual space is obtained, or may be when software is obtained.
Here, in the information update thread, processing to update the spatial information managed by spatial information managers 1201 and 1211 is performed.
The role of the information update thread is, for example, processing to update, based on the position and orientation of the VR goggles worn by the listener, the position and orientation of the listener's avatar positioned in the virtual space, updating of the positions of objects that have moved within the virtual space, and the like. Such processing is covered within a processing thread that launches relatively infrequently, at approximately tens of Hz.
Processing to update the information indicating a characteristic of a direct sound may be performed by such a processing thread that is created infrequently. The reason for this is that characteristics of direct sound vary less frequently than audio processing frames for audio output occur. This makes it possible to relatively reduce the computational load of the processing. Furthermore, when the information is updated at an unduly high frequency, there is a risk of pulsive noise being generated. Updating the information at a low frequency also makes it possible to avoid such a risk.
[Third Variation of Threshold Value Setting Method]
As another example of a threshold value setting method, threshold values may be set in accordance with computation resources (CPU capability, memory resources, PC performance, remaining level of battery, etc.) for processing reproduction of the virtual space. More specifically, sensor 1405 of audio signal processing device 1001 detects the amount of computation resources, and when the amount of computation resources is low, the threshold values are set to be high. Since consequently, the sound volume of a greater number of reflected sounds falls below the threshold values, the number of reflected sounds on which binaural processing is to be performed can be reduced, whereby the amount of computation can be reduced.
Alternatively, when the signal processing is performed by equipment that is driven by a storage battery, such as a smartphone or VR goggles, it is expected that priority is given to allowing processing to be performed for a longer duration, and computation resources are used economically. In such a case, it is not necessary to detect the amount or remaining level of computation resources, and the threshold values may be set to be high.
[Fourth Variation of Threshold Value Setting Method]
As another example of a threshold value setting method, by including a threshold value setter, not illustrated, in audio signal processing device 1001 or audio presentation device 1002, threshold values can be set by the manager of the virtual space or the listener.
For example, an “energy-saving mode”, in which there are few reflected sounds to be heard and the amount of computation is low, or a “high-performance mode”, in which there are many reflected sounds to be heard and the amount of computation is high, may be selectable by the listener to whom audio presentation device 1002 is equipped. Alternatively, the mode may be selectable by the manager who manages three-dimensional sound reproduction system 1000 or by the creator of the three-dimensional sound content. Furthermore, not the mode, but the threshold values or the threshold value data may be directly selectable.
[First Variation of Operations of Renderer]
FIG. 20 is a flowchart illustrating a first variation of operations of audio signal processing device 1001. FIG. 20 illustrates the processing performed mainly by renderer 1300 of audio signal processing device 1001. In this variation, sound volume compensation processing is added to the operations of renderer 1300.
For example, analyzer 1301 obtains data (the input signal) (S301). Next, analyzer 1301 analyzes the data (S302). Next, selector 1302 determines whether to select reflected sounds based on the analysis results (S303). Next, reproducer 1303 performs sound volume compensation processing based on the reflected sounds that were not selected (S304). Next, reproducer 1303 performs acoustic processing on the direct sounds and the reflected sounds (S305). Reproducer 1303 then outputs the direct sounds and the reflected sounds as audio (S306).
The above-described processes (S301 to S306) other than the sound volume compensation processing (S304) are processes that are shared with the other examples described above; thus, explanation thereof has been omitted.
The sound volume compensation processing is performed in accordance with the reflected sounds that were not selected in the selection processing. For example, due to not selecting a reflected sound in the selection processing, an absence emerges in the sound volume sensation. The sound volume compensation processing reduces the incongruity that accompanies this absence in the sound volume sensation. As an example of compensating the sound volume sensation, the following two methods are disclosed. Either of these two methods may be used.
First, a method in which the sound volume sensation is compensated for by raising the sound volume of a direct sound will be described. Reproducer 1303 raises the sound volume of the direct sound by the amount of the sound volume of a reflected sound that was not selected, and generates the direct sound. Accordingly, the sound volume sensation lost due to the reflected sound not being generated is compensated for.
At the time of raising the sound volume, reproducer 1303 may raise the sound volume of each frequency component in accordance with the frequency characteristics of the reflected sound. In order to make such processing possible, an attenuation rate of the sound volume attenuated by the reflection object may be assigned to each of predetermined frequency bands. This makes it possible to derive the frequency characteristics of the reflected sound.
Next, a method in which the sound volume sensation is compensated for by causing a reflected sound to be synthesized in a direct sound will be described. In this method, reproducer 1303 adds, to a direct sound, a reflected sound that was not selected and generates the direct sound to compensate for the sound volume sensation that results from the reflected sound not being generated. The sound volume (amplitude), frequency, delay, and the like of the reflected sound that was not selected are reflected in the generated direct sound.
In the case of the method for raising the sound volume of the direct sound, while the amount of computation for the compensation processing is extremely slight, only the sound volume is compensated for. In the case of the method of causing a reflected sound to be synthesized in a direct sound, the amount of computation for the compensation processing is large compared to the method of raising the sound volume of the direct sound, but the characteristics of the reflected sound are more accurately compensated for.
Since in both cases, only the direct sound is generated and the reflected sound is not generated, the total amount of computation is reduced. In particular, since the amount of computation required for binaural processing, which includes processing to implement a head-related transfer function (HRTF), is reduced, the total amount of computation is greatly reduced. The reason for this is that the amount of computation required for binaural processing is much greater than the amount of processing required for the above-described compensation processing.
Note that when the reason for a reflected sound not being selected is that the sound volume of the reflected sound is less than the masking threshold value, the sound volume sensation is not lost; thus, the reflected sound may be simply removed without performing compensation processing.
[Second Variation of Operations of Renderer]
FIG. 21 is a flowchart illustrating a second variation of operations of audio signal processing device 1001. FIG. 21 illustrates the processing performed mainly by renderer 1300 of audio signal processing device 1001. In this variation, left-right sound volume difference adjustment processing is added to the operations of renderer 1300.
For example, analyzer 1301 analyzes the input signal (S401). Next, analyzer 1301 detects the direction of arrival of sounds (S402). Next, selector 1302 adjusts the difference in sound volume between the sounds perceived by the left and right ears (S403). Furthermore, selector 1302 adjusts the difference in the arrival time periods (delay) between the sounds perceived by the left and right ears (S404). Selector 1302 determines whether to select reflected sounds based on information on the adjusted sounds (S405).
The above-described processes (S401 to S405) other than the left-right sound volume difference adjustment processing (S403) and the delay adjustment processing (S404) are processes that are shared with the other examples described above; thus, explanation thereof has been omitted.
FIG. 22 is a diagram illustrating an arrangement example of an avatar, a sound source object, and an obstacle object. For example, in a case in which the front direction of the listener is 0 degrees, when, as in FIG. 22, the polarities (for example, positive-negative) of the direction of arrival (θ) of the direct sound and the direction of arrival (γ) of the reflected sound (the direction (γ) of the reflected sound) are different, the sound volume difference that occurs between the ears is corrected.
Specifically, when the polarities of θ and γ are different, the ear which mainly (first) perceives the sound is different for each of the direct sound and the reflected sound. In this case, as the left-right sound volume difference adjustment processing (S403), selector 1302 adjusts the sound volume of the direct sound in accordance with the position of the ear that mainly perceives the reflected sound. For example, by multiplying the sound volume when the direct sound arrives at the listener by (1.0−0.3 sin(θ))(0≤θ≤180), selector 1302 causes attenuation of the sound volume when the direct sound arrives at the listener.
By calculating the sound volume ratio of the sound volume of the reflected sound to the sound volume of the direct sound, corrected as described above, and comparing the calculated sound volume ratio with threshold values, selector 1302 determines whether to select reflected sounds. Accordingly, the sound volume difference that occurs between the ears is corrected, the sound volume of direct sounds that affect reflected sounds is more accurately derived, and the determination of whether to select reflected sounds is more accurately performed.
Furthermore, in addition to the left-right sound volume difference adjustment processing (S403), selector 1302 may, as delay adjustment processing (S404), delay the direct sound arrival time period in accordance with the positions of the ears at which a reflected sound is perceived. Specifically, selector 1302 may delay the direct sound arrival time period by adding, to the direct sound arrival time period, (a (sin θ+θ)/c) ms (where a is the radius of the head and c is the speed of sound).
[Third Variation of Operations of Renderer]
A method for setting threshold values in accordance with directions of arrival will be described.
FIG. 23 is a flowchart illustrating yet another example of the selection processing. Description has been omitted for processes that are shared with the example in FIG. 14. In the example in FIG. 23, selector 1302 selects reflected sounds by using threshold values in accordance with directions of arrival.
Specifically, from the direct sound arrival path (pd), the reflected sound arrival path (pr), and avatar orientation information D1, each calculated by analyzer 1301, selector 1302 calculates the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (the direction (γ) of the reflected sound), each defined using the orientation of an avatar as reference. In other words, selector 1302 detects the direct sound arrival direction (θ) and the reflected sound arrival direction (γ) (S231). The orientation of the avatar corresponds to the orientation of the listener. Avatar orientation information D1 may be included in the input signal.
By using three indexes including the time difference (T), in addition to the direct sound arrival direction (θ) and the reflected sound arrival direction (γ), selector 1302 identifies, from a three-dimensional arrangement such as that illustrated in FIG. 15, the threshold values to be used in the selection processing (S232).
As an example, a method for setting threshold values to be used in selection processing when, as in FIG. 22, an avatar, a sound source object, and an obstacle object are arranged will be described.
Position information on the avatar, the sound source object, and the obstacle object, and avatar orientation information D1 are obtained from the input information. The direction (θ) of the direct sound and the direction (γ) of the sound image of the reflected sound when the orientation of the avatar is determined to be 0 degrees are calculated by using these items of position information and orientation information D1. In the case of FIG. 22, the direction (θ) of the direct sound is about 20 degrees, and the direction (γ) of the sound image of the reflected sound is about 265 degrees (−95 degrees).
Next, referencing the threshold value data stored in the three-dimensional arrangement illustrated in FIG. 15, a threshold value is identified from an arrangement domain that corresponds to the values of the two directions (θ) and (γ), and the value of the time difference (T) calculated by analyzer 1301. When there is no index that corresponds to the values of (θ), (γ), and (T) that were calculated, the threshold value corresponding to the index that is closest may be identified.
As another method, threshold values may be identified by performing processing such as interpolation or extrapolation, based on one or more threshold values that correspond to one or more indexes that are closest to the values of (θ), (γ), and (T) that were calculated. For example, a threshold value corresponding to (20 degrees, 265 degrees, T) may be identified based on the four threshold values corresponding to the four indexes of (0 degrees, 225 degrees, T), (0 degrees, 270 degrees, T), (45 degrees, 225 degrees, T), and (45 degrees, 270 degrees, T).
Selection processing based on the difference between the direct sound arrival direction angle (θ) and the reflected sound arrival direction angle (γ) will be described.
For example, as illustrated in FIG. 16, threshold value data having, as a two-dimensional index arrangement: the angular difference (φ) between the direct sound arrival direction (θ) and the reflected sound arrival direction (γ); and the time difference (T) may be pre-created and set. In this case, the angular difference (φ) and the time difference (T) are referenced in the selection processing. Alternatively, the angular difference (φ) between the angle (θ) of the direct sound arrival direction and the angle (γ) of the reflected sound arrival direction may be calculated in the selection processing, and the angular difference (φ) calculated may be used to identify the threshold value.
Alternatively, threshold value data having, as an index arrangement, a combination of the angular difference (φ), the direct sound arrival direction (θ), and the time difference (T), or a combination of the angular difference (φ), the reflected sound arrival direction (γ), and the time difference (T) may be set.
Alternatively, as illustrated in FIG. 15, threshold value data having, as a three-dimensional index arrangement, values of (θ), (γ), and (T) may be set.
[Fourth Variation of Operations of Renderer]
The processing performed by the above-described analyzer 1301, selector 1302, and reproducer 1303 may, for example, be performed as pipeline processing as described in PTL 3.
FIG. 24 is a block diagram illustrating a configuration example for renderer 1300 to perform pipeline processing.
Renderer 1300 in FIG. 24 includes reverberation processor 1311, early reflection processor 1312, distance attenuation processor 1313, selector 1314, generator 1315, and binaural processor 1316. These constituent elements may be configured as a plurality of the constituent elements of renderer 1300 illustrated in FIG. 7, or may be configured as at least a part of the plurality of constituent elements of audio signal processing device 1001 illustrated in FIG. 5.
Pipeline processing refers to dividing the processing for applying acoustic effects into a plurality of processes and executing each of the plurality of processes one by one in order. The plurality of processes include, for example, signal processing on the audio signal, generation of parameters used for signal processing, and the like.
Renderer 1300 may perform reverberation processing, early reflection processing, distance attenuation processing, binaural processing, and the like as pipeline processing. However, these types of processing are examples, and the pipeline processing may include processes other than these, or may not include some of these processes. For example, the pipeline processing may include diffraction processing and occlusion processing. Furthermore, for example, the reverberation processing may be omitted when unneeded.
Furthermore, each process may be expressed as a stage. Moreover, the audio signals of the reflected sounds and the like generated as the result of the processes may be expressed as rendering items. The plurality of stages and the order of these stages in the pipeline processing are not limited to the example illustrated in FIG. 24.
Here, the parameters (the arrival paths, the arrival time periods, and the sound volume ratios related to direct sounds and reflected sounds) used in the selection processing are calculated in one of the plurality of stages for generating the rendering items. In other words, the parameters used for selecting the reflected sounds are calculated as a part of the pipeline processing for generating the rendering items. Note that it is not necessary for all of the stages to be performed by renderer 1300. For example, a part of the stages may be omitted, or may be performed by an element other than renderer 1300.
The reverberation processing, the early reflection processing, the distance attenuation processing, the selection processing, the generation processing, and the binaural processing that may be included as stages in the pipeline processing will be described. In each stage, the metadata included in the input signal may be analyzed, and the parameters used for generating the reflected sounds may be calculated.
In the reverberation processing, reverberation processor 1311 generates an audio signal indicating reverberation sound or the parameters used in generating the audio signal. Reverberation sound is a sound that arrives at the listener as reverberation after the direct sound. As one example, the reverberation sound is a sound that arrives at the listener at a relatively late stage (for example, approximately 100 to 200 ms after the arrival of the direct sound) after the early reflected sound (to be described later) arrives at the listener, and after undergoing more reflections (for example, several tens of times) than the early reflected sound.
Reverberation processor 1311 refers to the audio signal and spatial information included in the input signal, and calculates reverberation sound by using, as a function for generating reverberation sound, a predetermined function prepared beforehand.
Reverberation processor 1311 may generate reverberation sound by applying a known reverberation generation method to the audio signal included in the input signal. One example of a known reverberation generation method is the Schroeder method, but the known reverberation generation method is not limited to the Schroeder method. Furthermore, reverberation processor 1311 uses the shape and acoustic characteristics of a sound reproduction space indicated by the spatial information when applying the known reverberation generation method. In this way, reverberation processor 1311 can calculate parameters for generating reverberation sound.
In the early reflection processing, early reflection processor 1312 calculates parameters for generating early reflection sounds based on the spatial information. The early reflected sound is reflected sound that arrives at the listener at a relatively early stage (for example, approximately several tens of ms after the arrival of the direct sound) after the direct sound from the sound source object arrives at the listener, and after undergoing one or more reflections.
Early reflection processor 1312 references, for example, the audio signal and metadata, and calculates the path, from reflection objects, of reflected sound that arrives at the listener after being reflected by the reflection objects. For example, in calculation of the path, the shape of the three-dimensional sound field (space), the size of the three-dimensional sound field, the positions of reflection objects such as structures, the reflectance of reflection objects, and the like may be used.
Early reflection processor 1312 may calculate the path of the direct sound. The information of said path may be used as a parameter for early reflection processor 1312 to generate the early reflected sound, and may be used as a parameter for selector 1314 to select reflected sounds.
In the distance attenuation processing, distance attenuation processor 1313 calculates the sound volume of the direct sound and the reflected sound that arrive at the listener, based on the lengths of the paths of the direct sound and the reflected sound. The sound volume of the direct sound and the reflected sound that arrive at the listener attenuate, with respect to the sound volume of the sound source, in proportion to the distance of the path to the listener (in inverse proportion to the distance). Thus, distance attenuation processor 1313 is able to calculate the sound volume of the direct sound by dividing the sound volume of the sound source by the length of the direct sound path, and is able to calculate the sound volume of the reflected sound by dividing the sound volume of the sound source by the length of the path of the reflected sound.
In the selection processing, selector 1314 selects the reflected sounds to be generated, based on the parameters calculated before the selection processing. One of the selection methods of the present disclosure may be used for selection of the reflected sounds to be generated.
The selection processing may be performed on all of the reflected sounds, or may be performed only on the reflected sounds having high evaluation values based on the evaluation processing, as described above. In other words, the reflected sounds having low evaluation values may be determined as not selected, without performing the selection processing. For example, reflected sounds for which the sound volume is extremely low may be considered to be reflected sounds having low evaluation values, and may be determined as not selected.
Furthermore, for example, the selection processing may be performed on all of the reflected sounds. Then, the evaluation values of the reflected sounds selected in the selection processing may be determined, and the reflected sounds having low determined evaluation values may be redetermined as not selected.
The selection processing and the evaluation processing may each be performed independently, or may be performed in combination with each other. When the selection processing and the evaluation processing are performed in combination with each other, either of the two processes may be performed first.
In the generation processing, generator 1315 generates direct sounds and reflected sounds. For example, generator 1315 generates direct sounds based on the direct sound arrival times and arrival time sound volume, from the audio signal included in the input signal. Furthermore, for each reflected sound selected in the selection processing, generator 1315 generates the reflected sound based on the reflected sound arrival time and the arrival time sound volume, from the audio signal included in the input signal.
In the binaural processing, binaural processor 1316 performs signal processing so that the audio signal of the direct sound is perceived as sound arriving at the listener from the direction of the sound source object. Furthermore, binaural processor 1316 performs signal processing so that the reflected sounds selected by selector 1314 are perceived as sounds arriving at the listener from the reflection object.
For example, based on the position and orientation of the listener in the sound space, binaural processor 1316 performs processing to apply an HRIR DB so that sound arrives at the listener from the position of the sound source object or the position of the obstacle object.
Note that Head-Related Impulse Response (HRIR) is the response characteristic when one impulse is generated. Specifically, HRIR is the response characteristic obtained by converting from an expression in the frequency domain to an expression in the time domain by Fourier transforming the head-related transfer function, in which the change in sound caused by surrounding objects including the auricle, the head, and the shoulders is expressed as a transfer function. The HRIR DB is a database including such information.
Furthermore, the position and orientation of the listener in the sound space are, for example, the position and orientation of a virtual listener in a virtual sound space. The position and orientation of the virtual listener in the virtual sound space may change in accordance with movement of the head of the listener. Furthermore, the position and orientation of the virtual listener in the virtual sound space may be determined based on information obtained from sensor 1405.
The program(s), spatial information, HRIR DB, threshold value data, other parameters, and/or the like used in the above-described processing are obtained from memory 1404 included in audio signal processing device 1001, or from outside of audio signal processing device 1001.
Furthermore, the pipeline processing may contain other processes. Moreover, renderer 1300 may contain a processor that is not illustrated, for performing another process included in the pipeline processing. For example, renderer 1300 may include a diffraction processor and an occlusion processor.
The diffraction processor executes processing to generate an audio signal indicating sound including diffracted sound caused by an obstacle object between the listener and the sound source object in a three-dimensional sound field (space). Diffracted sound is sound that when an obstacle object is present between the sound source object and the listener, arrives at the listener from the sound source object by going around the obstacle object.
The diffraction processor references, for example, the audio signal and metadata, and calculates the path by which diffracted sound arrives at the listener from the sound source object by detouring around the obstacle object, and generates diffracted sound based on the calculated path. In the calculation of the path, the sound source object in the three-dimensional sound field (space), the positions of the listener and the obstacle object, the shape and size of the obstacle object, and the like may be used.
When a sound source object is present on the other side of an obstacle object, the occlusion processor generates an audio signal for a sound that passes from the sound source object through the obstacle object and is audible therethrough, based on the spatial information and information on the material, etc. of the obstacle object.
[Example of Sound Source Object]
As described above, in the position information assigned to the sound source object, a “point” in the virtual space indicates the position of a sound source object. In other words, as described above, the sound source is defined as a “point sound source”.
On the other hand, a sound source in a virtual space may be defined as an object that has a length, size, shape, and the like, i.e., as a sound source that is not a point sound source, but a spatially extended sound source. In this case, the distance between the listener and the sound source, and the direction of arrival of the sound are not determined. Consequently, reflected sounds originating from such a sound source may be limited to being selected by selector 1302 without performing analysis by analyzer 1301, or regardless of the analysis result. By doing so, it is possible to avoid the sound quality degradation that might occur by not selecting the reflected sound.
Alternatively, a representative point such as the center of gravity of the object may be determined, and the processing of the present disclosure may be applied on the assumption that sound is generated from that representative point. In this case, the threshold value may be adjusted in accordance with information on the spatial extension of the sound source.
[Examples of Direct Sound and Reflected Sound]
For example, direct sound is sound that has not been reflected by a reflection object, and reflected sound is sound that has been reflected by a reflection object. Direct sound may be sound that has arrived at the listener from a sound source without being reflected by a reflection objection, and reflected sound may be sound that has arrived at the listener from a sound source due to being reflected by a reflection object.
Furthermore, each of direct sound and reflected sound are not limited to being sound that has arrived at the listener, and may each be sound that will arrive at the listener. For example, direct sound may be sound that has been outputted from a sound source, or to put it differently, a sound source sound.
FIG. 25 is a diagram illustrating transmission and diffraction of a sound. As illustrated in FIG. 25, a direct sound may not arrive at the listener due to the presence of an obstacle object between the sound source object and the listener. In this case, a sound that arrives at the listener after being emitted from the sound source object and passing through the obstacle object may be considered to be a direct sound. Furthermore, a sound that arrives at the listener after being emitted from the sound source object and diffracted by the obstacle object may be considered to be a reflected sound.
Furthermore, the two sounds compared in the selection processing are not limited to a direct sound and a reflected sound based on sound emitted from one sound source. For example, the selection of a sound may be performed by performing a comparison between two reflected sounds based on a sound emitted from one sound source. In this case, the direct sound in the present disclosure may be understood to be the sound that reaches the listener first, and the reflected sound in the present disclosure may be understood to be the sound that reaches the listener afterward.
[Example of Bitstream Structure]
The bitstream includes, for example, an audio signal and metadata. The audio signal is sound data in which sound is expressed, and indicates, e.g., information on the frequency and intensity of sound. Furthermore, metadata includes spatial information on the sound space, which is the space of the sound field.
For example, the spatial information is information on the space in which the listener who hears sound based on the audio signal is positioned. Specifically, the spatial information is information about a predetermined position (localization position) in the sound space (for example, a three-dimensional sound field) for localizing the sound image of the sound at that predetermined position, that is, for causing the listener to perceive the sound as arriving from a direction that corresponds to the predetermined position. The spatial information includes, for example, sound source object information and position information indicating the position of the listener.
The sound source object information is information on a sound source object that generates sound based on the audio signal. In other words, the sound source object information is information on an object (a sound source object) that reproduces the audio signal, and is information on a virtual sound source object located in a virtual sound space. Here, the virtual sound space may correspond to real-world space in which an object that generates sound is located, and the sound source object in the virtual sound space may correspond to an object that generates sound in a real-world space.
The sound source object information may indicate, for example, the position of the sound source object located in the sound space, the orientation of the sound source object, the directivity of the sound emitted by the sound source object, whether the sound source object belongs to an animate thing, whether the sound source object is a mobile body, and the like. For example, the audio signal is associated with one or more sound source objects indicated by the sound source object information.
The bitstream includes, for example, metadata (control information) and an audio signal.
The audio signal and metadata may be contained in a single bitstream or may be separately contained in a plurality of bitstreams. Furthermore, the audio signal and metadata may be contained in a single file or may be separately contained in a plurality of files.
The bitstream may exist for each sound source or may exist for each playback time. Even in a case in which bitstreams exist for each playback time, a plurality of bitstreams may be processed in parallel simultaneously.
Metadata may be assigned to each bitstream, or may be collectively assigned to a plurality of bitstreams as information for controlling the plurality of bitstreams. In this case, the plurality of bitstreams may share the metadata. Furthermore, the metadata may be assigned for each playback time.
When a plurality of bitstreams or a plurality of files exist, information indicating a relevant bitstream or a relevant file may be contained in one or more bitstreams or one or more files.
Alternatively, information indicating a relevant bitstream or a relevant file may be contained in each of all of the bitstreams or each of all of the files.
Here, the relevant bitstream or the relevant file is, for example, a bitstream or file that may be used simultaneously during acoustic processing. Furthermore, a bitstream or file that collectively describes the information indicating the relevant bitstream or the relevant file may be included.
Here, the information indicating the relevant bitstream or the relevant file may be, for example, an identifier indicating a relevant bitstream or a relevant file. Furthermore, the information indicating the relevant bitstream or the relevant file may be, for example, a file name indicating a relevant bitstream or a relevant file, a uniform resource locator (URL), a uniform resource identifier (URI), or the like.
In this case, an obtainer identifies and obtains a relevant bitstream or a relevant file based on the information indicating the relevant bitstream or the relevant file. Furthermore, the information indicating the relevant bitstream or the relevant file may be included in a bitstream or a file, and the information indicating the relevant bitstream or the relevant file may be included in a different bitstream or a different file.
Here, the file including the information indicating the relevant bitstream or the relevant file may be, for example, a control file such as a manifest file used in content distribution.
Note that the entire metadata or part of the metadata may be obtained from somewhere other than a bitstream of the audio signal. For example, either one of metadata for controlling an acoustic sound or metadata for controlling a video may be obtained from somewhere other than from a bitstream, or both may be obtained from somewhere other than from a bitstream.
Furthermore, the metadata for controlling a video may be included in the bitstream obtained by three-dimensional sound reproduction system 1000. In this case, three-dimensional sound reproduction system 1000 may output the metadata for controlling a video to a display device that displays images or a stereoscopic video reproduction device that reproduces stereoscopic videos.
[Examples of Information Included in Metadata]
The metadata may be information used for describing a scene expressed in the sound space. As used herein, the term “scene” refers to a collection of all elements that represent three-dimensional video and acoustic events in the sound space, which are modeled in three-dimensional sound reproduction system 1000 using metadata.
Thus, the metadata may include not only information for controlling acoustic processing, but also information for controlling video processing. The metadata may include only one among the information for controlling acoustic processing or the information for controlling video processing, or may include both.
Three-dimensional sound reproduction system 1000 generates virtual acoustic effects by performing acoustic processing on the audio signal using the metadata included in the bitstream and additionally obtained interactive listener position information. Early reflection processing, obstacle processing, diffraction processing, occlusion processing, and reverberation processing may be performed as acoustic effects, and other acoustic processing may be performed using the metadata. For example, an acoustic effect such as a distance decay effect, localization, or a Doppler effect may be added.
In addition, information for switching between on and off of all or one or more of the acoustic effects, and priority information regarding a plurality of processes for the acoustic effects may be added to the metadata.
As an example, the metadata includes information about a sound space including a sound source object and an obstacle object and information about a localization position for localizing the sound image at a predetermined position in the sound space (that is, causing the listener to perceive the sound as arriving from a predetermined direction).
Here, an obstacle object is an object that can influence a sound emitted by a sound source object and perceived by the listener, by, for example, blocking or reflecting the sound between the sound source object and the listener. The obstacle object can include an animal or a movable body such as a machine, in addition to a stationary object. The animal may be a person or the like.
Furthermore, when a plurality of sound source objects are present in a sound space, another sound source object may be an obstacle object for a certain sound source object. In other words, non-sound-emitting objects such as building materials or inanimate objects, and sound source objects that emit sound can both be obstacle objects.
The metadata includes information indicating all or part of the shape of the sound space, the shapes and positions of obstacle objects in the sound space, the shapes and positions of sound source objects in the sound space, and the position and orientation of the listener in the sound space.
The sound space may be either a closed space or an open space. Furthermore, the metadata may include information indicating the reflectance of each obstacle object that can reflect sound in the sound space. For example, the floor, walls, ceiling, and the like constituting the boundaries of the sound space can be included in the obstacle objects.
The reflectance is an energy ratio between a reflected sound and an incident sound, and may be set for each sound frequency band. Of course, the reflectance may be uniformly set, irrespective of the sound frequency band. Note that when the sound space is an open space, for example, parameters such as a uniformly set attenuation rate, diffracted sound, and early reflected sound may be used.
The metadata may include information other than reflectance as a parameter with regard to an obstacle object or a sound source object. For example, the metadata may include information on the material of an object as a parameter related to both of a sound source object and a non-sound-emitting object. Specifically, the metadata may include information such as the diffusivity, transmittance, and sound absorption rate.
For example, information on a sound source object may include information indicating, for example, sound volume, radiation characteristics (directivity), a reproduction condition, the number and types of sound sources of one object, and a sound source region of an object. The reproduction condition may determine whether a sound is, for example, a sound that is continuously being emitted or is emitted at an event. The sound source region of an object may be determined by the relative relationship between the position of the listener and the position of the object, or may be determined using the object as a reference.
For example, when the sound source region is determined by the relative relationship between the position of the listener and the position of the object, it is possible to cause the listener to perceive sound E from the right side of the object and sound F from the left side of the object, the right side and the left side being as seen from the listener.
Furthermore, when the sound source region is determined using the object as a reference, it is possible to fix what sound is emitted from what region of the object, using the object as a reference. For example, it is possible, when the listener sees the object from the front, to cause the listener to perceive a high sound from the right side of the object and a low sound from the left side of the object. Furthermore, it is possible, when the listener sees the object from the rear, to cause the listener to perceive a low sound from the right side of the object and a high sound from the left side of the object.
Metadata related to the space may include the time period until early reflected sound, the reverberation time period, the ratio of direct sound to diffuse sound, and the like. When the ratio between a direct sound and a diffuse sound is zero, the listener can be caused to perceive only the direct sound.
BRIEF SUMMARY
Here, the present embodiment is briefly summarized.
When the relationship between a direct sound and a reflected sound is analyzed and the direct sound is the leading sound and the reflected sound is the lagging sound, if the relationship is such that the precedence effect occurs, i.e., if the reflected sound falls below the echo detection limit, the reflected sound is not perceived; thus, the auditory impact on the listener is small, even if the reflected sound is removed.
FIG. 26 is a diagram illustrating an example of the positional relationship between a listener and an obstacle object, according to the present embodiment. FIG. 27 is a diagram illustrating another example of the positional relationship between a listener and an obstacle object, according to the present embodiment. It should be noted that FIG. 26 has the same positional relationship as that illustrated in FIG. 9, and FIG. 27 has the same positional relationship as that shown in FIG. 10. Furthermore, FIG. 28 is an example of an echo detection limit threshold value according to the present embodiment. It should be noted that the echo detection limit threshold value shown in FIG. 28 is an example of the threshold value data shown in FIG. 12C and the like.
For example, comparing the positional relationship in FIG. 26 with the positional relationship in FIG. 27, the sound volume of the reflected sound heard by the listener at the positional relationship in FIG. 26 is lower than the sound volume of the reflected sound heard by the listener at the positional relationship in FIG. 27. This is because the path length until the reflected sound arrives in the positional relationship in FIG. 26 is longer than the path length until the reflected sound arrives in the positional relationship in FIG. 27.
Thus, when making the determination based solely on the sound volume of the reflected sound, the reflected sound illustrated in FIG. 26 has less auditory impact than the reflected sound illustrated in FIG. 27. However, comparing the arrival time of the reflected sound at the listening position illustrated in FIG. 26 with the arrival time of the reflected sound at the listening position illustrated in FIG. 27, the reflected sound arrives later in the case illustrated in FIG. 26.
Therefore, when the determination is made in terms of the echo detection limit, as illustrated in FIG. 28, the reflected sound illustrated in FIG. 27 is below the echo detection limit and is thus not perceived by the listener as a reflected sound, while the reflected sound illustrated in FIG. 26 is above the echo detection limit and is thus perceived by the listener as a reflected sound.
In the present embodiment, the amount of computation involved in processing reflected sounds is reduced by utilizing this to determine the auditory importance of reflected sounds, and not reproducing reflected sounds that are unimportant.
The above description is a brief summary of the present embodiment.
Here, attention is directed to sound volume ratio L and the threshold value.
The threshold value (echo detection limit threshold value) that is compared to sound volume ratio L is based on the auditory precedence effect. Therefore, as described above, when sound volume ratio L is calculated based solely on physical characteristics, the result of selecting whether to reproduce the reflected sound (the selection result) may not align with the listener's actual perception.
That is, without taking the auditory sensitivity of the listener into consideration, whether to output an audio signal indicating the reflected sound (more specifically, an output signal based on that audio signal) is selected, and there are cases in which such an output signal is output and the listener hears the sound (the reflected sound) represented by that output signal. In such cases, the listener hears a sound that differs from his/her own auditory perception, leading to a sense of incongruity.
Therefore, the following is a more detailed description of an audio signal processing method that can appropriately reduce the amount of computation and the computational load in a sound space, while taking auditory sensitivity into consideration.
Embodiment 2
Embodiment 2 is described below. The description below is centered on the points of difference from Embodiment 1, and descriptions of points in common are omitted or simplified.
[Configuration of Renderer]
First, the configuration of renderer 2300 according to the present embodiment is described. FIG. 29 is a block diagram illustrating a configuration example of renderer 2300 according to the present embodiment.
Renderer 2300 includes analyzer 2301, selector 2302, and reproducer 2303. It should be noted that as described above, the audio signal processing device according to the present embodiment is an example of a decoding device. The decoding device includes a decoder, and the decoder includes renderer 2300. In other words, it can be stated that the audio signal processing device according to the present embodiment includes analyzer 2301, selector 2302, and reproducer 2303. Renderer 2300 applies acoustic processing to sound data included in the input signal, and outputs the result.
Similarly to Embodiment 1, the input signal includes, for example, spatial information, sensor information, and sound data. The spatial information also includes physical information such as reflection coefficients, transmission coefficients, and diffraction coefficients of non-emitting objects (obstacle objects).
It should be noted that in the present embodiment, mainly reflected sound, which is an example of indirect sound, is used for description, but the same processing is performed even if indirect sound other than reflected sound is used. Furthermore, examples of indirect sound include reflected sound, diffracted sound, and the like.
Analyzer 2301 may be able to perform all of the processing or some of the processing performed by analyzer 1301 according to Embodiment 1.
Similarly to analyzer 1301 according to Embodiment 1, analyzer 2301 performs analysis on the audio signal included in the input signal, as well as the spatial information received from spatial information managers 1201 and 1211. Analyzer 2301 thus calculates the information necessary to generate direct sound and reflected sound with reproducer 2303, as well as the information necessary to select whether to generate reflected sound. The method by which analyzer 2301 calculates these items of information is as described in Embodiment 1.
Analyzer 2301 also performs analysis processing on the input signal, as performed by analyzer 1301 according to Embodiment 1, in S101 of FIG. 8. In other words, analyzer 2301 analyzes the input signal input to the audio signal processing device according to the present embodiment to detect direct sound and reflected sound that may be generated in the sound space.
When such direct sound and reflected sound are detected, analyzer 2301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data.
More specifically, analyzer 2301 creates an audio signal indicating reflected sound and an audio signal indicating direct sound based on: the position information of the sound source objects, the position information of the non-sound emitting objects (obstacle objects), and the position information and physical information of the listener that are included in the spatial information; and the sound data.
Specifically, analyzer 2301 creates an audio signal generated within the virtual space, based on the spatial information and the sound data, and then assigns attribute information indicating an attribute identifying the audio signal to the created audio signal, thereby creating an audio signal containing the attribute information. An audio signal including attribute information is created for each sound generated in the virtual space. The attribute includes information indicating whether the sound indicated by the audio signal is a direct sound or a reflected sound (indirect sound). As an example, in the present embodiment, the attribute is information indicating whether the sound indicated by the relevant audio signal is direct sound or reflected sound. Furthermore, the attribute information may include information necessary to radiate the audio signal into the sound space, such as, for example, gain information, gain characteristic information for each frequency bandwidth, position information, directivity information, and the like. In other words, the relevant necessary information may be retained in the attribute information. Furthermore, attribute information may be tied to the audio signal as metadata. The gain characteristics for each frequency bandwidth of the audio signal included in the attribute information may be identified based on, for example, the spatial information included in the input information. Information indicating frequency characteristics that indicate auditory sensitivity may be identified based on, for example, the spatial information included in the input information, especially as information tied to the avatar of the listener.
It should be noted that for simplicity, an audio signal for which the attribute is information indicating reflected sound (indirect sound) may be described as an audio signal indicating reflected sound (indirect sound), and an audio signal for which the attribute is information indicating direct sound may be described as an audio signal indicating direct sound.
Sound that directly reaches the listener's head from one sound source is direct sound, and sound that reaches the listener's head after being output from that one sound source and then reflected by a non-emitting object or diffracted by a non-emitting object is indirect sound (reflected sound or diffracted sound).
It should be noted that in the present embodiment, analyzer 2301 creates an audio signal indicating a reflected sound (indirect sound) and an audio signal indicating the direct sound associated with that reflected sound (indirect sound).
Furthermore, a direct sound associated with an indirect sound means a direct sound that originates from the same sound source as that indirect sound. An indirect sound associated with a direct sound means an indirect sound that originates from the same sound source as that direct sound. Furthermore, a reflected sound is a sound resulting from the direct sound associated with that reflected sound being reflected by a reflector.
An audio signal for which the attribute is information indicating a reflected sound (indirect sound) includes information indicating an audio signal for the direct sound associated with that reflected sound (indirect sound).
Analyzer 2301 may cause the created audio signal to be stored in memory included in analyzer 2301. Furthermore, analyzer 2301 also creates a plurality of audio signals and causes the plurality of audio signals to be stored in the memory.
Furthermore, as in Embodiment 1, analyzer 2301 may calculate, for each of the direct sound and reflected sound, values related to: the path until arriving at the listening position; the time period taken until arrival; the sound volume at arrival; and the like. Similarly, analyzer 2301 may calculate information indicating the relationship between the direct sound and the reflected sound, such as, for example, a value related to the time difference between when the direct sound arrives and when the reflected sound arrives (the time difference between the direct sound and the reflected sound), and the like.
It should be noted that the reflected sound arrival time sound volume (Ir) and the direct sound arrival time sound volume (Id) are examples of the arrival time sound volume and the like. The direct sound arrival time sound volume (Id) refers to the sound volume of a direct sound at the time of arrival at the listening position, which is the location at which the listener is present in a virtual space. In other words, the direct sound arrival time sound volume (Id) is the sound volume of the direct sound at the listening position. The reflected sound arrival time sound volume (Ir) refers to the sound volume of a reflected sound, which is an example of an indirect sound, at the time of arrival at the listening position. In other words, the reflected sound arrival time sound volume (Ir) is the sound volume of the indirect sound (the sound volume of the reflected sound) at the listening position.
In the present embodiment, each of the audio signal whose attribute is information indicating reflected sound and the audio signal whose attribute is information indicating direct sound may include information indicating the sound volume, at the listening position, of the sound indicated by the audio signal. In other words, in the present embodiment, the audio signal whose attribute is information indicating reflected sound includes information indicating the reflected sound arrival time sound volume (Ir) as the sound volume of the indirect sound (sound volume of the reflected sound). Similarly, the audio signal whose attribute is information indicating direct sound includes information indicating the direct sound arrival time sound volume (Id) as the sound volume of the direct sound.
It should be noted that the audio signal whose attribute indicates reflected sound may also include information indicating the sound volume of the indirect sound (sound volume of the reflected sound) and the sound volume of the direct sound associated with that reflected sound. Similarly, the audio signal whose attribute is information indicating direct sound may include information indicating the sound volume of the direct sound and the sound volume of the indirect sound (sound volume of the reflected sound) associated with that direct sound.
Selector 2302 may be able to perform all of the processing or some of the processing performed by selector 1302 according to Embodiment 1. Furthermore, selector 2302 determines whether the output signal based on the audio signal created by analyzer 2301 is to be output (reproduced) by reproducer 2303. That is, selector 2302 first specifies one audio signal from among the plurality of audio signals created by analyzer 2301 (for example, the audio signals indicating reflected sounds), and then selects whether reproducer 2303 is to generate and output an output signal based on the one audio signal specified.
Selector 2302 has obtainer 2302a, first calculator 2302b, second calculator 2302c, and selection processor 2302d.
Obtainer 2302a obtains an audio signal that includes attribute information and was created by analyzer 2301 and stored in the memory of analyzer 2301. Obtainer 2302a obtains, for example, an audio signal indicating a reflected sound (indirect sound) and an audio signal indicating a direct sound associated with the reflected sound (indirect sound). Furthermore, obtainer 2302a also obtains values related to the time difference between the direct sound and the reflected sound calculated by analyzer 2301.
First calculator 2302b calculates the first sound volume based on the audio signal that indicates the indirect sound and was obtained by obtainer 2302a. Here, first calculator 2302b calculates the first sound volume that is based on the sound volume of the indirect sound, which is the sound volume of the indirect sound indicated by the audio signal at the time at which the indirect sound arrives at the listening position that is the position at which the listener is present.
More specifically, first calculator 2302b calculates the first sound volume that is based on the sound volume of the reflected sound (that is, the reflected sound arrival time sound volume (Ir)), which is the sound volume of the reflected sound indicated by the audio signal at the time at which the reflected sound arrives at the listening position that is the position at which the listener is present.
Second calculator 2302c calculates the second sound volume based on the audio signal that indicates the direct sound and was obtained by obtainer 2302a. Here, second calculator 2302c calculates the second sound volume that is based on the sound volume of the direct sound (in other words, the direct sound arrival time sound volume (Id)), which is the sound volume of the direct sound indicated by the audio signal at the time at which the direct sound arrives at the listening position that is the position at which the listener is present.
It should be noted that in the present embodiment, the second sound volume is a sound volume different from the sound volume of the direct sound, but the sound volume of the direct sound may be used as the second sound volume, as-is.
Furthermore, selection processor 2302d selects whether reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound (indirect sound), based on the sound volume ratio between the second sound volume calculated and the first sound volume calculated.
When selection processor 2302d selects that reproducer 2303 is to output the output signal, selector 2302 outputs the audio signal obtained to reproducer 2303.
Reproducer 2303 may be able to perform all of the processing or some of the processing performed by reproducer 1303 according to Embodiment 1. Furthermore, reproducer 2303 obtains the audio signal output from selector 2302 and outputs an output signal that is based on the audio signal obtained.
Reproducer 2303 generates and outputs the output signal by performing binaural filtering processing and/or the like on the audio signal obtained. Binaural filtering processing is realized, for example, by subjecting the audio signal obtained to processing using a head-related transfer function.
Furthermore, reproducer 2303 may synthesize and output the audio signal that indicates the direct sound and was obtained by obtainer 2302a and the output signal generated.
Furthermore, reproducer 2303 may also generate and output an output signal by performing both binaural filter processing and diffusion filter processing on the audio signal output from selector 2302. The diffusion filter processing is, for example, processing that improves the realism of indirect sound by diffusing the reflected sound (indirect sound) indicated by the audio signal in the audio signal obtained. Furthermore, the diffusion filter processing is processing in which a filter is used to simulate the auditory intensity of the sound diffusion indicated by the audio signal obtained (i.e., to simulate the auditory intensity of the sound diffusion as perceived by the listener). Finite impulse filters and/or infinite impulse filters are used in the diffusion filter processing.
Hereinafter, an example of the operation of the audio signal processing method performed by the audio signal processing device according to the present embodiment (more specifically, renderer 2300) is described.
[Operation Example of Renderer]
FIG. 30 is a flowchart illustrating an operation example of the audio signal processing device according to the present embodiment. FIG. 30 illustrates the processing performed mainly by renderer 2300 included in the audio signal processing device according to the present embodiment. It should be noted that here, descriptions of common points with FIG. 8 according to Embodiment 1 are omitted or simplified.
First, analyzer 2301 performs analysis processing to analyze the input signal (S101a). More specifically, analyzer 2301 analyzes the input signal to detect direct sound and reflected sound that may be generated in the sound space. When such direct sound and reflected sound are detected, analyzer 2301 creates audio signals including attribute information, i.e., an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. Analyzer 2301 causes the created audio signal to be stored in the memory of analyzer 2301.
Analyzer 2301 analyzes the input signal and calculates, for each of the direct sound and reflected sound, values related to, e.g., the path until arrival at the listening position, the time taken to arrive, and the sound volume at the time of arrival, as well as a value related to the time difference between the direct sound and the reflected sound.
First, analyzer 2301 calculates the characteristics of each of the direct sound indicated by the audio signal created and the reflected sound indicated by the audio signal created. Specifically, the arrival time period and the arrival time sound volume when each of the direct sound and the reflected sound arrive at the listener (listening position) are calculated. It should be noted that the method shown in Embodiment 1 may be used as the method for calculating these arrival time periods and arrival time sound volumes.
It should be noted that as described above, the audio signal indicating reflected sound includes information indicating the reflected sound arrival time sound volume (Ir) as the sound volume of the reflected sound, and the audio signal indicating direct sound includes information indicating the direct sound arrival time sound volume (Id) as the sound volume of the direct sound.
Then, analyzer 2301 calculates the time difference (T) between the direct sound and the reflected sound (the time difference (T) between when the direct sound arrives and when the indirect sound arrives). It should be noted that the method shown in Embodiment 1 may be used as the method for calculating the time difference (T). Unlike step S101 according to Embodiment 1, the sound volume ratio (L) need not be calculated in step S101a.
Selector 2302 (more specifically, selection processor 2302d) performs the selection of reflected sounds (selection processing) (S102a). In other words, selector 2302 selects whether reproducer 2303 is to reproduce an output signal that is based on an audio signal that indicates a reflected sound and was created by analyzer 2301.
First, obtainer 2302a obtains an audio signal that includes attribute information and was created by analyzer 2301 and stored in the memory. Obtainer 2302a obtains, for example, at least one of an audio signal indicating a reflected sound (indirect sound) or an audio signal indicating a direct sound associated with the reflected sound (indirect sound). Here, obtainer 2302a obtains both. Furthermore, obtainer 2302a also obtains a value related to the time difference (T) between the direct sound and the reflected sound calculated by analyzer 2301.
As described above, the reflected sound and the direct sound associated with that reflected sound originate from the same sound source.
First calculator 2302b calculates the first sound volume that is based on the sound volume of the reflected sound, based on the audio signal that indicates the reflected sound and was obtained by obtainer 2302a.
Second calculator 2302c obtains the second sound volume that is based on the sound volume of the direct sound, based on the audio signal that indicates the direct sound and was obtained by obtainer 2302a.
Furthermore, selection processor 2302d selects whether reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound (indirect sound), based on: the sound volume ratio between the second sound volume calculated and the first sound volume calculated; and the time difference (T) between the direct sound and the reflected sound (indirect sound).
More specifically, selection processor 2302d selects that reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound (indirect sound), when the sound volume ratio is greater than or equal to the first threshold value determined according to the time difference (T) between the direct sound and the reflected sound (indirect sound).
The sound volume ratio is a value obtained by dividing the first sound volume, which is based on the reflected sound arrival time sound volume (Ir), by the second sound volume, which is based on the direct sound arrival time sound volume (Id).
Furthermore, the time difference (T) is the time difference (T) between: the direct sound associated with the reflected sound indicated by the audio signal obtained; and the reflected sound indicated by the audio signal obtained, and is the time difference (T) between the direct sound and the reflected sound, calculated by analyzer 2301 in step S101a. As described in Embodiment 1, the time difference (T) between a direct sound and a reflected sound is, for example, the time difference between the direct sound arrival time period (arrival time) and the reflected sound arrival time period (arrival time), but is not limited thereto.
The first threshold value is a value determined according to the time difference (T) between the direct sound associated with the reflected sound (indirect sound) and the reflected sound (indirect sound); in other words, the threshold value is a value dependent on the time difference (T) and is the value indicated by the threshold value data of Embodiment 1. The threshold value data is, for example, a graph having a horizontal axis that indicates the time difference (T) between a direct sound and reflected sounds and a vertical axis that indicates the sound volume ratios of reflected sounds to a direct sound, and is expressed as a threshold value (first threshold value) that demarcates whether each reflected sound is perceived.
More specifically, the threshold value data indicating the first threshold value is the data shown in FIG. 11 to FIG. 13 and the like.
Furthermore, in step S102a, obtainer 2302a obtains the gain characteristic of each predetermined frequency bandwidth related to the indirect sound (reflected sound) and the frequency characteristic indicating the auditory sensitivity. Similarly, in step S102a, obtainer 2302a obtains the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
The gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the frequency characteristic indicating the auditory sensitivity, and the gain characteristic of each predetermined frequency bandwidth related to the direct sound are stored, for example, in the memory of analyzer 2301. Obtainer 2302a obtains, from the memory, the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the frequency characteristic indicating the auditory sensitivity, and the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
It should be noted that the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the frequency characteristic indicating the auditory sensitivity, and the gain characteristic of each predetermined frequency bandwidth related to the direct sound may be obtained by obtainer 2302a via a communication line or the like.
Selector 2302 performs the selection processing as described above.
Next, the selection processing, particularly the calculation of the first sound volume and the second sound volume, is described in greater detail with reference to FIG. 31.
FIG. 31 is a flowchart illustrating an operation example of the selection processing according to the present embodiment. It should be noted that here, descriptions of common points with FIG. 14 according to Embodiment 1 are omitted or simplified.
First, selector 2302 specifies a reflected sound detected by analyzer 2301 (S201). In other words, obtainer 2302a of selector 2302 specifies the audio signal that includes the attribute information and was created by analyzer 2301 and stored in the memory, and obtains the audio signal specified. For example, obtainer 2302a specifies an audio signal indicating the reflected sound and obtains the audio signal specified. At this time, obtainer 2302a may also obtain an audio signal indicating the direct sound associated with that reflected sound (indirect sound).
Then, selector 2302 calculates the first sound volume that is based on the sound volume of the reflected sound (S210).
More specifically, first calculator 2302b of selector 2302 calculates the first sound volume that is based on the sound volume of the reflected sound, based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal indicating the reflected sound.
Furthermore, selector 2302 calculates the second sound volume that is based on the sound volume of the direct sound (S220).
More specifically, second calculator 2302c of selector 2302 calculates the second sound volume that is based on the sound volume of the direct sound, based on: a second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound. Here, second calculator 2302c calculates the second sound volume that is based on the sound volume of the direct sound, based on the second correction characteristic and the audio signal indicating the direct sound. It should be noted that when the audio signal indicating the reflected sound includes information indicating the sound volume of the direct sound associated with the reflected sound, second calculator 2302c may calculate the second sound volume that is based on the sound volume of the direct sound, based on the second correction characteristic and the audio signal indicating the reflected sound.
Below, the calculations of the first sound volume and the second sound volume are described. The first sound volume and the second sound volume are calculated using Formula 5 and Formula 6 below.
First Sound Volume=GlobalGain_R*EqualiserGain_R (Formula 5)
Second sound volume=GlobalGain_D*EqualiserGain_D (Formula 6)
It should be noted that GlobalGain_R is the gain of the signal of the reflected sound across all frequency bands, and EqualiserGain_R is the representative value of the gain of the equalizer applied to the reflected sound for each frequency band. GlobalGain_D is the gain of the signal of the direct sound across all frequency bands, and EqualiserGain_D is the representative value of the gain of the equalizer applied to the direct sound for each frequency band.
The gain of the signal across all frequency bands may be calculated using the sound volume of the sound source of the sound and the distance from the sound source to the listener, as described in Embodiment 1. More specifically, the gain of the signal of the reflected sound across all frequency bands corresponds to the sound volume of the indirect sound (the sound volume of the reflected sound); in other words, it corresponds to the arrival time sound volume (Ir) included in the audio signal indicating the reflected sound. Furthermore, the gain of the signal of the direct sound across all frequency bands corresponds to the sound volume of the direct sound; in other words, it corresponds to the direct sound arrival time sound volume (Id) included in the audio signal indicating the direct sound.
It should be noted that when obtainer 2302a obtains only the audio signal indicating the reflected sound, the sound volume of the direct sound included in the audio signal indicating the reflected sound may be used as the gain of the signal of the direct sound across all frequency bands.
As described above, the representative value indicated by EqualiserGain_R is the representative value of the gain of the equalizer applied to the reflected sound for each frequency band. The gain of the equalizer applied to the reflected sound for each frequency band corresponds to the first correction characteristic described above. That is, EqualiserGain_R is the representative value of the first correction characteristic obtained by correcting, using the frequency characteristic indicating the auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the reflected sound.
For simplicity, the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound) may be referred to as the gain characteristic related to the reflected sound (indirect sound).
The gain characteristic related to the reflected sound (indirect sound) may be the gain, for each band, of an equalizer that realizes an adjustment in the frequency characteristic resulting from the fact that, for example, when a sound wave strikes an object (a non-sound-emitting object), the reflectance varies for each frequency component depending on the surface shape, material, or hardness of that object.
FIG. 32 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth related to a reflected sound according to the present embodiment. The gain characteristic related to reflected sound is indicated by the reflection coefficient, as shown in FIG. 32, but is not limited thereto.
FIG. 32 illustrates examples of gain characteristics corresponding to each of three types of wall surfaces. FIG. 32 shows the gain characteristic of each of a wall surface having a hard surface with numerous irregularities, a wall surface having a hard surface with few irregularities, and a wall surface having a soft surface. The gain characteristic corresponding to the object (non-sound-emitting object) generating the reflected sound is employed. It should be noted that the gain characteristics illustrated in FIG. 32 are merely examples.
Furthermore, while a ⅓-octave band is used here as the predetermined frequency bandwidth, this is not intended to be limiting.
The frequency characteristic indicating the auditory sensitivity is a frequency characteristic indicating the sound volume sensitivity of the listener, and for example, an A-weighting characteristic may be used. The A-weighting characteristic is a frequency-weighting characteristic that takes human hearing into consideration. FIG. 33A is a diagram illustrating a table showing a frequency characteristic indicating auditory sensitivity according to the present embodiment. FIG. 33B is a diagram illustrating the frequency characteristic sensitivity according to the present indicating the auditory embodiment. FIG. 33B is a diagram illustrating the frequency characteristic with the vertical axis expressed in decibels (dB). It should be noted that the frequency characteristic (A-weighting characteristic) indicating the auditory sensitivity illustrated in FIG. 33B displays a value for each ⅓-octave band, but this is not intended to be limiting.
Further, EqualiserGain_R is calculated as follows. Here, description is provided with reference to FIG. 34.
FIG. 34 is a diagram illustrating the gain characteristic related to the reflected sound according to the present embodiment, the frequency characteristic indicating the auditory sensitivity (A-weighting characteristic), and the first correction characteristic. FIG. 34 shows a value for each ⅓-octave band.
Using the above-described frequency characteristic (A-weighting characteristic) indicating the auditory sensitivity, the gain characteristic related to the reflected sound is corrected. Specifically, the value of the first correction characteristic corresponding to a given band is calculated by multiplying the value of the frequency characteristic indicating the auditory sensitivity for each ⅓-octave band, by the value of the gain characteristic related to the reflected sound corresponding to that band. In FIG. 34, since the vertical axis is represented on a logarithmic scale (dB), the value of the first correction characteristic is calculated by, for each frequency band, adding the value of the frequency characteristic indicating the auditory sensitivity and the value of the gain characteristic related to the reflected sound. Furthermore, when the vertical axis in FIG. 34 is expressed as a linear axis (scaling factor), it goes without saying that the calculation may be performed using multiplication.
The representative value of the first correction characteristic calculated in this manner is denoted by EqualiserGain_R. It should be noted that the representative value of the first correction characteristic may be, e.g., the average value, the maximum value, or the minimum value of the values of the first correction characteristic, or may be, e.g., the average value, the maximum value, or the minimum value of the first correction characteristic within a predetermined frequency band.
Furthermore, EqualiserGain_D is described.
As described above, the representative value indicated by EqualiserGain_D is the representative value of the gain of the equalizer applied to the direct sound for each frequency band. The gain of the equalizer applied to the direct sound for each frequency band corresponds to the second correction characteristic described above. That is, EqualiserGain_D is the representative value of the second correction characteristic obtained by correcting, using the frequency characteristic indicating the auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
For simplicity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound may be referred to as the gain characteristic related to the direct sound.
The gain characteristic related to the direct sound may, for example, be the gain of the equalizer for each band, indicating the gain of frequency components due to the state of the space through which the sound propagates. More specifically, the gain characteristic related to the direct sound may be the gain, for each band, of the equalizer realizing a frequency characteristic exhibiting a tendency such that higher frequency components experience greater gain attenuation due to factors such as the temperature or humidity of the space through which the sound propagates, or the density of fine particles like dust or pollen. FIG. 35 is a diagram illustrating a gain characteristic of each predetermined frequency bandwidth of a direct sound according to the present embodiment. The gain characteristic related to direct sound is not limited to that shown in FIG. 35.
As shown in FIG. 35, for the gain characteristic related to direct sound, a value is shown for each predetermined frequency bandwidth, that is, for each ⅓-octave band, but this is not intended to be limiting.
It should be noted that the frequency characteristic indicating the auditory sensitivity used to calculate EqualiserGain_D may also be the A-weighting characteristic illustrated in FIG. 33B.
Further, EqualiserGain_D is calculated as follows. Here, a description is provided with reference to FIG. 36.
FIG. 36 is a diagram illustrating the gain characteristic related to the direct sound, the frequency characteristic indicating the auditory sensitivity (A-weighting characteristic), and the second correction characteristic, according to the present embodiment. FIG. 36 illustrates the value for each ⅓-octave band.
Using the above-described frequency characteristic that indicates the auditory sensitivity (A-weighting characteristic), the gain characteristic related to the direct sound is corrected. Specifically, the value of the second correction characteristic corresponding to a given band is calculated by multiplying the value of the frequency characteristic indicating the auditory sensitivity for each ⅓-octave band, by the value of the gain characteristic related to the direct sound corresponding to that band. In FIG. 36, since the vertical axis is represented on a logarithmic scale (dB), the value of the second correction characteristic is calculated by, for each frequency band, adding the value of the frequency characteristic indicating the auditory sensitivity and the value of the gain characteristic related to the direct sound. Furthermore, when the vertical axis in FIG. 36 is expressed as a linear axis (scaling factor), it goes without saying that the calculation may be performed using multiplication.
The representative value of the second correction characteristic calculated in this manner is denoted by EqualiserGain_D. It should be noted that the representative value of the second correction characteristic may be, e.g., the average value, the maximum value, or the minimum value of the values of the second correction characteristic, or may be, e.g., the average value, the maximum value, or the minimum value of the second correction characteristic within a predetermined frequency band.
The first sound volume and the second sound volume are calculated using EqualiserGain_D and EqualiserGain_R calculated in this way, according to Formula 5 and Formula 6 described above.
That is, the first sound volume is calculated by multiplying the reflected sound arrival time sound volume (Ir) (the sound volume of the reflected sound) included in the audio signal indicating the reflected sound by the calculated EqualiserGain_R calculated. Similarly, the second sound volume is calculated by multiplying the direct sound arrival time sound volume (Id) (the sound volume of the direct sound) included in the audio signal indicating the direct sound by EqualiserGain_D calculated.
It should be noted that in the present embodiment, the A-weighting characteristic was used as the frequency characteristic indicating the auditory sensitivity, that is, the frequency characteristic indicating the sound volume sensitivity of the listener; however, this is not intended to be limiting. For example, the inverse characteristic of the equal loudness contour, the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position, or the like may be used as the frequency characteristic indicating the sound volume sensitivity of the listener.
FIG. 37 is a diagram illustrating the inverse characteristic of the equal loudness contour, which is another, first example of the frequency characteristic indicating the sound volume sensitivity of the listener according to the present embodiment. In FIG. 33B, which shows the A-weighting characteristic, the A-weighting characteristic exhibits a convex upward characteristic. Similarly, the inverse characteristic of the equal loudness contour exhibits a convex upward characteristic. Furthermore, the method of using the inverse characteristic of the equal loudness contour in the above-described correction may be the same as described above.
FIG. 38 is a diagram illustrating the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position, which is another, second example of the frequency characteristic indicating the sound volume sensitivity of the listener according to the present embodiment. Furthermore, the method of using the inverse frequency characteristic of the minimum audible angle of the sound source position in the above-described correction may be the same as described above.
Thus, in the present embodiment, the frequency characteristic indicating the sound volume sensitivity of the listener can be utilized as the frequency characteristic indicating the auditory sensitivity. Furthermore, as the frequency characteristic indicating the sound volume sensitivity, a frequency characteristic based on the inverse of the equal loudness contour, or the inverse characteristic of the frequency characteristic of the minimum audible angle of the sound source position can be utilized.
Once more, a description is provided with reference to FIG. 31. Next, selection processor 2302d calculates the sound volume ratio between the second sound volume calculated and the first sound volume calculated (S202a).
Then, selection processor 2302d detects the time difference (T) between the direct sound and the reflected sound (S203). The time difference (T) has already been calculated by analyzer 2301 in step S101a. For example, data indicating the time difference (T) is stored in the memory of analyzer 2301, and selector 2302 detects the time difference (T) by obtaining this data.
Furthermore, selection processor 2302d identifies the first threshold value corresponding to the time difference (T), using the threshold value data (S204). Then, selection processor 2302d determines whether or not the sound volume ratio calculated is greater than or equal to the first threshold value (S205a).
When the sound volume ratio is greater than or equal to the first threshold value (“Yes” in S205a), selection processor 2302d selects the reflected sound as a reflected sound to be generated (S206). Specifically, in this case, selection processor 2302d selects that reproducer 2303 is to reproduce an output signal that is based on the audio signal that indicates the reflected sound and was created by analyzer 2301.
When the sound volume ratio is less than the first threshold value (“No” in S205a), selection processor 2302d skips selecting the reflected sound as a reflected sound to be generated (S207). Specifically, in this case, selection processor 2302d selects that reproducer 2303 is not to reproduce the output signal that is based on the audio signal that indicates the reflected sound and was created by analyzer 2301, thereby determining that the reflected sound is a reflected sound that is not to be generated, i.e., a reflected sound to be culled.
Subsequently, selection processor 2302d determines whether there are any unspecified reflected sounds (S208). That is, selection processor 2302d determines whether any of the plurality of audio signals created by analyzer 2301 have not undergone selection processing. If there are any unspecified reflected sounds (“Yes” in S208), selection processor 2302d repeats the above-described processing (S201 to S207, S210, and S220). If there are no unspecified reflected sounds (“No” in S208), selection processor 2302d ends the processing.
The selection processing is thus performed. In step S206, when selection processor 2302d selects that reproducer 2303 is to reproduce the output signal that is based on the audio signal indicating the reflected sound, selector 2302 outputs the audio signal to reproducer 2303.
Once more, a description is provided with reference to FIG. 30. Reproducer 2303 obtains the audio signal output from selector 2302 and outputs an output signal that is based on the audio signal (S103a). Here, reproducer 2303 synthesizes and outputs the audio signal that indicates the direct sound and was obtained by obtainer 2302a, and the sound signal generated (the audio signal indicating the reflected sound).
Thus, in step S205a, when the sound volume ratio is greater than or equal to the first threshold value, that is, when selection processor 2302d selects that reproducer 2303 is to reproduce an output signal that is based on the audio signal indicating the reflected sound, reproducer 2303 outputs the output signal that is based on that audio signal.
It should be noted that when the sound volume ratio is less than the first threshold value in step S205a, that is, when selection processor 2302d does not select that reproducer 2303 is to reproduce an output signal that is based on the audio signal indicating the reflected sound, reproducer 2303 skips outputting the output signal that is based on that audio signal. In such cases, reproducer 2303 does not output an output signal that is based on the audio signal, thereby reducing the amount of computation and the computational load.
As described above, the audio signal processing method according to the present embodiment is an audio signal processing method executed by an audio signal processing device (renderer 2300), and includes an obtaining step, a first calculating step, a second calculating step, a selection processing step, and a reproducing step.
In the obtaining step, an audio signal that includes attribute information identifying the attribute of the audio signal is obtained. The attribute includes information indicating an indirect sound (e.g., a reflected sound). In the first calculating step, a first sound volume is calculated. The first sound volume is based on the sound volume of the indirect sound (the sound volume of the reflected sound) at a time at which the indirect sound arrives at a listening position that is a position at which a listener is present. The calculating of the first sound volume is based on: a first correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the indirect sound; and the audio signal obtained. In the second calculating step, a second sound volume is calculated. The second sound volume is based on a sound volume of a direct sound, associated with the indirect sound, at a time at which the direct sound arrives at the listening position. In the selection processing step, whether to output an output signal that is based on the audio signal obtained is selected, based on: a sound volume ratio between the second sound volume calculated and the first sound volume calculated; and a time difference between when the direct sound arrives and when the indirect sound arrives. In the reproducing step, the output signal is output, when outputting the output signal is selected.
Whether an output signal based on the audio signal indicating the indirect sound is output is thus selected, based on: the sound volume ratio between the second sound volume that is based on the sound volume of the direct sound and the first sound volume that is based on the sound volume of the indirect sound; and the time difference (T). In other words, whether to output the output signal that is based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Furthermore, the first sound volume is calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the first sound volume calculated is used; and the time difference (T). That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
Furthermore, in the audio signal processing method according to the present embodiment, in the second calculating step, the second sound volume is calculated based on the second correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the direct sound.
The second sound volume is thus calculated taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal is selected based on: the sound volume ratio in which the second sound volume calculated is used; and the time difference. That is, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into further consideration.
Furthermore, in the audio signal processing method according to the present embodiment, the frequency characteristic indicating the auditory sensitivity is the frequency characteristic indicating the sound volume sensitivity of the listener, and the frequency characteristic indicating the sound volume sensitivity is the A-weighting characteristic.
This makes it possible to realize an audio signal processing method that enables using the frequency characteristic indicating the sound volume sensitivity of the listener as the frequency characteristic indicating the auditory sensitivity, and enables using the A-weighting characteristic as the frequency characteristic indicating the sound volume sensitivity.
Embodiment 3
Embodiment 3 is described below. The description below is centered on the points of difference from Embodiments 1 and 2, and descriptions of points in common are omitted or simplified.
In Embodiment 2, the output signal was selected to be output when the sound volume ratio between the second sound volume and the first sound volume was greater than or equal to the first threshold value. Embodiment 3 is similar to Embodiment 2 in the respect that the output signal is selected to be output when the sound volume ratio (L) is greater than or equal to the first threshold value. However, the method for determining the first threshold value in Embodiment 3 differs from that in Embodiment 2.
[Configuration of Renderer]
First, the configuration of renderer 3300 according to the present embodiment is described. FIG. 39 is a block diagram illustrating a configuration example of renderer 3300 according to the present embodiment.
Renderer 3300 includes analyzer 3301, selector 3302, and reproducer 2303.
Analyzer 3301 differs from analyzer 2301 according to Embodiment 2 in that analyzer 3301 calculates, e.g., a value related to the sound volume ratio (L) between a direct sound and a reflected sound at the listening position.
That is, analyzer 3301 detects direct sound and reflected sound that may be generated in the sound space, and when such direct sound and reflected sound are detected, analyzer 3301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on spatial information and sound data.
In the present embodiment, a direct sound and a reflected sound that is a sound resulting from the direct sound being reflected by a reflector (e.g., a non-sound-emitting object) are detected by analyzer 3301. The direct sound is also a sound associated with the reflected sound. Analyzer 3301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound associated with the reflected sound.
Furthermore, as in Embodiments 1 and 2, analyzer 3301 may calculate, for each of the direct sound and reflected sound, values related to: the path until arriving at the listening position; the time period taken until arrival; the sound volume at arrival; and the like. Analyzer 3301 then calculates values representing information indicating the relationship between the direct sound and the reflected sound, such as, for example, a value related to the time difference (T) between the direct sound and the reflected sound (the time difference (T) between when the direct sound arrives and when the reflected sound arrives), a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, and the like. The method by which analyzer 3301 calculates these items of information is as described in Embodiment 1.
It should be noted that as described in Embodiment 2, the sound volume of the reflected sound (the sound volume of the indirect sound) is the reflected sound arrival time sound volume (Ir), and the sound volume of the direct sound is the direct sound arrival time sound volume (Id). In other words, the sound volume ratio (L), calculated by analyzer 3301, between the direct sound and the reflected sound at the listening position is the sound volume ratio (L) between the sound volume of the direct sound and the sound volume of the reflected sound.
It should be noted that analyzer 3301 may be able to perform all of the processing or some of the processing performed by analyzer 1301 according to Embodiment 1.
Selector 3302 has obtainer 3302a and selection processor 3302d. Selector 3302 may be able to perform all of the processing or some of the processing performed by selector 1302 according to Embodiment 1.
Obtainer 3302a obtains the audio signal indicating the reflected sound and the audio signal indicating the direct sound associated with the reflected sound, both created by analyzer 3301. Furthermore, obtainer 3302a obtains, e.g., a value related to the time difference (T) between the direct sound and the reflected sound calculated by analyzer 3301 and a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, each calculated by analyzer 3301.
Selection processor 3302d selects whether reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound, based on: the sound volume ratio (L) and the time difference (T) obtained by obtainer 3302a; and the reflection coefficient characteristic amount.
For example, selection processor 3302d performs selection processing as described in step S102 of FIG. 8 in Embodiment 1 and step S102a of FIG. 30 in Embodiment 2. Selection processor 3302d selects that reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound, when the sound volume ratio (L) calculated by analyzer 3301 is greater than or equal to the first threshold value determined according to the time difference (T) between the direct sound and the reflected sound.
Furthermore, the time difference (T) is the time difference (T) between: the direct sound associated with the reflected sound indicated by the audio signal obtained; and the reflected sound indicated by the audio signal obtained, and is the time difference (T) between the direct sound and the reflected sound calculated by analyzer 3301. As described in Embodiment 1, the time difference (T) between a direct sound and a reflected sound is, for example, the time difference between the direct sound arrival time period (arrival time) and the reflected sound arrival time period (arrival time), but is not limited thereto.
The first threshold value is a value determined according to the time difference (T) between the direct sound associated with the reflected sound and the reflected sound; in other words, the threshold value is a value dependent on the time difference (T) and is the value indicated by the threshold value data of Embodiment 1.
Moreover, the reflection coefficient characteristic amount is a characteristic amount determined based on the reflection coefficient of the reflector (e.g., a non-sound-emitting object). The reflection coefficient characteristic amount is a characteristic amount indicating the degree of flatness of the frequency characteristic of the reflection coefficient, such as that illustrated in FIG. 32. More specifically, the reflection coefficient characteristic amount is a characteristic amount indicating the degree of flatness of the frequency spectrum of the gain characteristic indicated by the reflection coefficient. For the wall surface having a hard surface with numerous irregularities and the wall surface having a hard surface with few irregularities, each shown in FIG. 32, the degree of flatness of the frequency characteristic of the reflection coefficient is high, and the reflection coefficient characteristic amount is large. For the wall surface having a soft surface illustrated in FIG. 32, the degree of flatness of the frequency characteristic of the reflection coefficient is low, and the reflection coefficient characteristic amount is small.
Furthermore, the reflection coefficient characteristic amount can also be described as a characteristic amount indicating the degree of variation in the attenuation amount, as indicated by the reflection coefficients shown in FIG. 32.
The reflection coefficient characteristic amount may be obtained by obtainer 3302a via a communication line or the like. Furthermore, the reflection coefficient characteristic amount may be stored in the memory included in analyzer 3301, and obtainer 3302a may obtain the reflection coefficient characteristic amount stored in the memory.
When selection processor 3302d performs the selection processing, the reflection coefficient characteristic amount is used as follows.
In the present embodiment, the first threshold value is changed according to the magnitude of the reflection coefficient characteristic amount. For example, when the reflection coefficient characteristic amount is large, the first threshold value is changed to be higher, and when the reflection coefficient characteristic amount is small, the first threshold value is changed to be lower.
FIG. 40 is a diagram illustrating the impact of the reflection coefficient characteristic amount on the first threshold value, according to the present embodiment. FIG. 40 illustrates an example of the echo detection limit threshold value (first threshold value), similar to FIG. 28. When the reflection coefficient characteristic amount is large, the first threshold value becomes higher, shifting such that the first threshold value increases across the entire range of the horizontal axis shown in FIG. 40, for example. When the reflection coefficient characteristic amount is small, the first threshold value becomes lower, shifting such that the first threshold value becomes lower across the entire range of the horizontal axis, as shown in FIG. 40, for example.
Here, the precedence effect and the reflection coefficient characteristic amount are examined.
The technique described in Embodiment 1 and the like, described above, utilizes the precedence effect. The precedence effect is said to occur when the frequency spectrum of a leading sound (for example, a direct sound) approximates that of a lagging sound (for example, a reflected sound). In other words, when the frequency spectrum of the leading sound does not approximate that of the lagging sound, it is considered that the precedence effect will not occur.
Incidentally, since reflected sound is sound resulting from a direct sound being reflected by a reflector (e.g., a non-sound-emitting object), the frequency spectrum of the lagging sound (reflected sound) depends on the reflection coefficient of the reflector, more specifically, on the reflection coefficient characteristic amount. Therefore, whether the frequency spectrum of the direct sound approximates the frequency spectrum of the reflected sound varies depending on the reflection coefficient characteristic amount.
For example, when the reflection coefficient of the reflector corresponds to the reflection coefficients of the wall surface having a hard surface and numerous irregularities and the wall surface having a hard surface and few irregularities, each shown in FIG. 32, the reflection coefficient characteristic amount is large. Consequently, the frequency spectrum of the leading sound approximates the frequency spectrum of the lagging sound, making the precedence effect more likely to occur. In this case, the first threshold value is changed to be higher, that is, selection is performed such that the output signal that is based on the audio signal is less likely to be output. This is because the auditory value of reflected sound decreases as the precedence effect is more likely to occur.
Furthermore, when the reflection coefficient of the reflector is that of the wall surface having a soft surface shown in FIG. 32, the reflection coefficient characteristic amount is small. Consequently, the frequency spectrum of the leading sound does not approximate the frequency spectrum of the lagging sound, making the precedence effect less likely to occur. In this case, the first threshold value is changed to be lower, that is, selection is performed such that the output signal based on the audio signal is more likely to be output.
Selecting whether to output the output signal based on the reflection coefficient characteristic amount is thus equivalent to selecting whether to output the output signal while taking the precedence effect into consideration. Since the precedence effect is an example of the auditory sensitivity characteristic, the audio signal processing method according to the present embodiment is able to appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
As described above, selection processor 3302d selects whether reproducer 2303 is to output an output signal based on the obtained audio signal indicating the reflected sound, based on: the sound volume ratio (L) and the time difference (T) obtained by obtainer 3302a; and the reflection coefficient characteristic amount.
When selection processor 3302d selects that reproducer 2303 is to output the output signal, selector 3302 outputs the audio signal obtained to reproducer 2303.
Reproducer 2303 obtains the audio signal output from selector 3302 and outputs an output signal that is based on the audio signal obtained.
Hereinafter, an example of the operation of the audio signal processing method performed by the audio signal processing device according to the present embodiment (more specifically, renderer 3300) is described.
[Operation Example of Renderer]
FIG. 41 is a flowchart illustrating an operation example of the audio signal processing device according to the present embodiment. FIG. 41 illustrates the processing performed mainly by renderer 3300 included in the audio signal processing device according to the present embodiment.
First, analyzer 3301 performs analysis processing to analyze the input signal (S101b). More specifically, analyzer 3301 analyzes the input signal to detect direct sound and reflected sound that may be generated in the sound space. When such direct sound and reflected sound are detected, analyzer 3301 creates audio signals including attribute information, i.e., an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. Analyzer 3301 causes the created audio signal to be stored in the memory of analyzer 3301. Furthermore, analyzer 3301 analyzes the input signal and calculates, for each of the direct sound and reflected sound, values related to, e.g., the path until arrival at the listening position, the time taken to arrive, and the sound volume at arrival; a value related to the time difference (T) between the direct sound and the reflected sound; and a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position.
That is, analyzer 3301 calculates the sound volume ratio (L), which is the ratio between the direct sound arrival time sound volume (Id) and the reflected sound arrival time sound volume (Ir), and the time difference (T) between the direct sound and the reflected sound. The sound volume ratio (L) is the sound volume ratio (L) between the direct sound and the reflected sound at the listening position. It should be noted that the methods shown in Embodiment 1 may be used as the methods for calculating the sound volume ratio (L) and the time difference (T).
It should be noted that unlike step S101a according to Embodiment 2, in step S101b according to the present embodiment, the sound volume ratio (L) is calculated.
Selector 3302 (more specifically, selection processor 3302d) performs the selection of reflected sounds (selection processing) (S102b). In other words, selector 3302 selects whether reproducer 2303 is to reproduce an output signal based on the audio signal that indicates the reflected sound and was created by analyzer 3301.
First, obtainer 3302a obtains an audio signal that includes attribute information and was created by analyzer 3301 and stored in the memory. Obtainer 3302a obtains, for example, an audio signal indicating a reflected sound and an audio signal indicating a direct sound associated with the reflected sound. Furthermore, obtainer 3302a obtains, e.g., a value related to the time difference (T) between the direct sound and the reflected sound calculated by analyzer 3301 and a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, each calculated by analyzer 3301. Furthermore, obtainer 3302a obtains the reflection coefficient characteristic amount stored in the memory of analyzer 3301.
Selection processor 3302d selects that reproducer 2303 is to output an output signal that is based on the obtained audio signal indicating the reflected sound, when the sound volume ratio (L) obtained is greater than or equal to the first threshold value.
The first threshold value is a value determined based on the time difference (T) between the direct sound and the reflected sound, and determined based on the reflection coefficient characteristic amount obtained. That is, selection processor 3302d determines the first threshold value based on the time difference (T) between the direct sound and the reflected sound, and the reflection coefficient characteristic amount.
As described above, selection processor 3302d selects whether reproducer 2303 is to output an output signal that is based on the audio signal obtained, based on the sound volume ratio (L) between the sound volume of the reflected sound and the sound volume of the direct sound, the time difference (T) between the direct sound and the reflected sound, and the reflection coefficient characteristic amount.
The selection processing is thus performed. When selection processor 3302d selects that reproducer 2303 is to reproduce the output signal that is based on the audio signal indicating the reflected sound, selector 3302 outputs the audio signal to reproducer 2303.
Then, reproducer 2303 obtains the audio signal output from selector 3302 and outputs an output signal that is based on the audio signal (S103b). Here, for example, reproducer 2303 synthesizes and outputs the audio signal indicating the direct sound obtained by obtainer 3302a and the sound signal generated (the audio signal indicating the reflected sound).
As described above, the audio signal processing method according to the present embodiment is an audio signal processing method executed by an audio signal processing device, and includes an obtaining step, a selection processing step, and a reproducing step.
In the obtaining step, an audio signal including attribute information identifying an attribute of the audio signal is obtained. The attribute includes information indicating a reflected sound that is a sound resulting from a direct sound being reflected by a reflector. That is, this attribute includes information indicating reflected sound. In the selection processing step, whether to output an output signal that is based on the audio signal obtained is selected, based on: a sound volume ratio between: a sound volume of the reflected sound, indicated by the audio signal obtained, at a time at which the reflected sound arrives at a listening position that is a position at which a listener is present; and a sound volume of the direct sound at a time at which the direct sound arrives at the listening position; a time difference between when the direct sound arrives and when the reflected sound arrives; and a reflection coefficient characteristic amount determined based on a reflection coefficient of the reflector. In the reproducing step, the output signal is output, when outputting the output signal is selected.
Whether to output the output signal that is based on the audio signal indicating the reflected sound is thus selected based on the above-described sound volume ratio and the above-described time difference. In other words, whether to output the output signal that is based on the audio signal is appropriately selected. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load.
Here, attention is directed to the precedence effect. The precedence effect is said to occur when the frequency spectrum of a leading sound (for example, a direct sound) approximates the frequency spectrum of a lagging sound (for example, a reflected sound). The frequency spectrum of reflected sound varies according to the reflection coefficient of the reflector. Therefore, whether the frequency spectrum of a direct sound approximates the frequency spectrum of a reflected sound varies in accordance with the reflection coefficient characteristic amount.
Selecting whether to output the output signal based on the reflection coefficient characteristic amount is thus equivalent to selecting whether to output the output signal while taking the precedence effect into consideration. Since the precedence effect is an example of an auditory sensitivity characteristic, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking the auditory sensitivity into consideration.
Furthermore, in the present embodiment, the reflection coefficient characteristic amount indicates the degree of flatness of the frequency characteristic of the reflection coefficient.
This makes it possible to realize an audio signal processing method that enables using the reflection coefficient characteristic amount indicating the degree of flatness in the frequency characteristics of the reflection coefficient.
Embodiment 4
Embodiment 4 is described below. The description below is centered on the points of difference from Embodiment 2, and descriptions of points in common are omitted or simplified.
In Embodiment 2, the output signal was selected to be output when the sound volume ratio between the second sound volume and the first sound volume was greater than or equal to the first threshold value. Embodiment 4 differs from Embodiment 2 in that whether to output the output signal is selected based on a predetermined sound volume corresponding to the sound volume of a sound at a time at which the sound arrives at the listener.
It should be noted that in Embodiment 2 and the like, indirect sound (reflected sound) and direct sound were distinguished. However, in the present embodiment, when distinguishing between indirect sound (reflected sound) and direct sound is unnecessary, both indirect sound (reflected sound) and direct sound may simply be referred to as “sound”.
[Configuration of Renderer]
First, the configuration of renderer 4300 according to the present embodiment is described. FIG. 42 is a block diagram illustrating a configuration example of renderer 4300 according to the present embodiment.
Renderer 4300 includes analyzer 4301, selector 4302, and reproducer 2303.
Analyzer 4301 detects sounds that may be generated in the sound space. When such sound is detected, analyzer 4301 creates an audio signal indicating that sound, based on spatial information and sound data.
More specifically, analyzer 4301 detects direct sound and reflected sound that may be generated in the sound space, and when such direct sound and reflected sound are detected, analyzer 4301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. In other words, the audio signal indicating the sound is either an audio signal indicating reflected sound or an audio signal indicating direct sound.
In the present embodiment, a direct sound and a reflected sound that is a sound resulting from the direct sound being reflected by a reflector (e.g., a non-sound-emitting object) are detected by analyzer 4301. Thus, analyzer 4301 creates an audio signal indicating the reflected sound and an audio signal indicating the direct sound associated with the reflected sound.
Furthermore, analyzer 4301 may calculate values related to the path the sound takes until arriving at the listening position, the time the sound takes to arrive, the arrival time sound volume, and the like. That is, as in Embodiments 1 and 2, for each of the direct sound and reflected sound, values related to the following may be calculated: the path until arriving at the listening position; the time period taken until arrival; the sound volume at arrival; and the like.
It should be noted that in the present embodiment, analyzer 4301 may not calculate values representing information indicating the relationship between the direct sound and the reflected sound, such as, for example, a value related to the time difference (T) between the direct sound and the reflected sound (the time difference (T) between when the direct sound arrives and when the reflected sound arrives), a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, and the like.
The audio signal according to the present embodiment includes information indicating the sound volume of the sound indicated by the audio signal, at the time at which the sound arrives at the listening position. That is, similar to Embodiment 2, an audio signal whose attribute is information indicating reflected sound includes information indicating the reflected sound arrival time sound volume (Ir), as the sound volume of the reflected sound at the time at which the reflected sound arrives at the listening position. An audio signal whose attribute is information indicating direct sound includes information indicating the direct sound arrival time sound volume (Id), as the sound volume of the direct sound at the time at which the direct sound arrives at the listening position.
It should be noted that analyzer 4301 may be able to perform all of the processing or some of the processing performed by analyzer 1301 according to Embodiment 1.
Selector 4302 has obtainer 4302a, third calculator 4302e, and selection processor 4302d.
Obtainer 4302a obtains audio signals that indicate sounds and were created by analyzer 4301. That is, obtainer 4302a obtains an audio signal indicating a reflected sound and an audio signal indicating a direct sound associated with the reflected sound.
It should be noted that as described above, the audio signals indicating sounds include information indicating the sound volume at the time at which the sound represented by that audio signal arrives at the listening position. That is, the audio signal indicating the reflected sound includes information representing the sound volume of the reflected sound (the reflected sound arrival time sound volume (Ir)), and the audio signal indicating the direct sound includes information representing the sound volume of the direct sound (the direct sound arrival time sound volume (Id)).
Third calculator 4302e calculates a predetermined sound volume based on the sound volume of the sound indicated by the audio signal at a time at which the sound arrives at the listening position, based on the audio signal indicating the sound. More specifically, third calculator 4302e calculates the predetermined sound volume based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal obtained by obtainer 4302a; and the audio signal obtained.
In the present embodiment, the audio signal indicating the sound is either an audio signal indicating reflected sound or an audio signal indicating direct sound.
When the audio signal indicating the sound is an audio signal indicating reflected sound, third calculator 4302e performs the following processing to calculate the predetermined sound volume.
In this case, the gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal is the same as the gain characteristic of each predetermined frequency bandwidth related to the reflected sound, described in Embodiment 2. Similarly, the frequency characteristic indicating the auditory sensitivity is, for example, an A-weighted characteristic. Further, the predetermined correction characteristic is the first correction characteristic described in Embodiment 2. Then, third calculator 4302e calculates the first sound volume as the predetermined sound volume by the same method as that in Embodiment 2, based on the predetermined correction characteristic (first correction characteristic) and the audio signal indicating the reflected sound. In other words, in this case, the predetermined sound volume is the first sound volume.
Furthermore, when the audio signal indicating the sound is an audio signal indicating direct sound, third calculator 4302e performs the following processing to calculate the predetermined sound volume.
In this case, the gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal is the same as the gain characteristic of each predetermined frequency bandwidth related to the direct sound, described in Embodiment 2. Similarly, the frequency characteristic indicating the auditory sensitivity is, for example, an A-weighted characteristic. Further, the predetermined correction characteristic is the second correction characteristic described in Embodiment 2. Then, third calculator 4302e calculates the second sound volume as the predetermined sound volume by the same method as that in Embodiment 2, based on the predetermined correction characteristic (second correction characteristic) and the audio signal indicating the direct sound. In other words, in this case, the predetermined sound volume is the second sound volume.
Selection processor 4302d selects whether reproducer 2303 is to output an output signal that is based on the audio signal obtained, based on the predetermined sound volume calculated by third calculator 4302e.
Selection processor 4302d selects that reproducer 2303 is to output an output signal that is based on the audio signal obtained, when the predetermined sound volume calculated is greater than or equal to the second threshold value.
The second threshold value, unlike the first threshold value, is a value independent of the time difference (T) between the direct sound associated with the reflected sound, and the reflected sound, and is a fixed value. The second threshold value is a value related to the sound volume obtained; in other words, the second threshold value is a value related to the amplitude value. Furthermore, the second threshold value indicates the sound volume of the boundary demarcating whether a sound is perceivable to the listener, and is a threshold value for determining a sound having a lower sound volume than the threshold value to be a sound that is not to be reproduced.
FIG. 43 is a graph illustrating threshold value data indicating the second threshold value according to the present embodiment. For example, the second threshold value is −70 dB. Since the predetermined sound volume (the second sound volume) of the audio signal indicating the direct sound in FIG. 43 is greater than or equal to the second threshold value, the output signal that is based on the audio signal is selected to be output. Furthermore, since the predetermined sound volume (the first sound volume) of the audio signal indicating the reflected sound in FIG. 43 is less than the second threshold value, the output signal that is based on that audio signal is selected not to be output.
It should be noted that selector 4302 may be able to perform all of the processing or some of the processing performed by selector 1302 according to Embodiment 1.
Hereinafter, an example of the operation of the audio signal processing method performed by the audio signal processing device according to the present embodiment (more specifically, renderer 4300) is described.
[Operation Example of Renderer]
FIG. 44 is a flowchart illustrating an operation example of the audio signal processing device according to the present embodiment. FIG. 44 illustrates the processing performed mainly by renderer 4300 included in the audio signal processing device according to the present embodiment.
First, analyzer 4301 performs analysis processing to analyze the input signal (S101c). More specifically, analyzer 4301 analyzes the input signal to detect sounds (direct sound and reflected sound) that may be generated in the sound space. When such sounds are detected, analyzer 4301 creates audio signals indicating the sounds, more specifically an audio signal indicating the reflected sound and an audio signal indicating the direct sound, based on the spatial information and the sound data. Analyzer 4301 causes the audio signal created to be stored in the memory of analyzer 4301.
Furthermore, analyzer 4301 analyzes the input signal to calculate, for each sound (the direct sound and the reflected sound), values related to the path until arrival at the listening position, the time taken until arrival, the arrival time sound volume, and the like.
It should be noted that in step S101c according to the present embodiment, it is not necessary to calculate values such as a value related to the time difference (T) between the direct sound and the reflected sound, a value related to the sound volume ratio (L) between the direct sound and the reflected sound at the listening position, and the like.
Selector 4302 (more specifically, selection processor 4302d) performs the selection of sounds (selection processing) (S102c). In other words, selector 4302 selects whether reproducer 2303 is to reproduce an output signal that is based on the audio signal that indicates the sound and was created by analyzer 4301.
First, obtainer 4302a obtains the audio signals (the audio signal indicating the reflected sound and the audio signal indicating the direct sound) that include attribute information and were created by analyzer 4301 and stored in the memory.
Then, third calculator 4302e calculates the predetermined sound volume based on the predetermined correction characteristic obtained by correcting, using the frequency characteristic indicating the auditory sensitivity, the gain characteristic of each predetermined frequency bandwidth related to the sound indicated by the audio signal obtained by obtainer 4302a; and the audio signal obtained.
When the audio signal indicating the sound is an audio signal indicating reflected sound, the predetermined sound volume is the first sound volume. Furthermore, when the audio signal indicating the sound is an audio signal indicating direct sound, the predetermined sound volume is the second sound volume.
Next, selection processor 4302d selects that reproducer 2303 is to output an output signal based on the audio signal obtained, when the predetermined sound volume calculated by third calculator 4302e is greater than or equal to the second threshold value.
Selector 4302 performs the selection processing as described above.
The selection processing is described in greater detail with reference to FIG. 45.
FIG. 45 is a flowchart illustrating an operation example of the selection processing according to the present embodiment.
First, selector 4302 specifies a sound detected by analyzer 4301 (S201c). In other words, obtainer 4302a of selector 4302 specifies the audio signal created by analyzer 4301 and stored in the memory, and obtains the audio signal specified.
Then, third calculator 4302e calculates the predetermined sound volume (the first sound volume or the second sound volume) (S230).
Next, selection processor 4302d determines whether or not the predetermined sound volume calculated is greater than or equal to the second threshold value (S205c).
When the predetermined sound volume is greater than or equal to the second threshold value (“Yes” in S205c), selection processor 4302d selects the sound indicated by the audio signal obtained as a sound to be generated (S206c). Specifically, in this case, selection processor 4302d selects that reproducer 2303 is to reproduce an output signal that is based on the audio signal that indicates the sound and was created by analyzer 4301.
When the predetermined sound volume is lower than the second threshold value (“No” in S205c), selection processor 4302d skips selecting the sound indicated by the audio signal obtained as a sound to be generated (S207c). Specifically, in this case, selection processor 4302d selects that reproducer 2303 is not to reproduce the output signal that is based on the audio signal that indicates the sound and was created by analyzer 4301, thereby determining that the sound is a sound that is not to be generated, i.e., a sound to be culled.
Subsequently, selection processor 4302d determines whether there are any unspecified sounds (S208c). That is, selection processor 4302d determines whether any of the plurality of audio signals created by analyzer 4301 have not undergone selection processing. If there are any unspecified sounds (“Yes” in S208c), selection processor 4302d repeats the above-described processing (S201c to S207c and S230). If there are no unspecified reflected sounds (“No” in S208c), selection processor 4302d ends the processing.
The selection processing is thus performed. In step S206c, when selection processor 4302d selects that reproducer 2303 is to reproduce the output signal based on the audio signal indicating the sound, selector 4302 outputs the audio signal to reproducer 2303.
Then, reproducer 2303 obtains the audio signal output from selector 4302 and outputs an output signal that is based on the audio signal (S103c).
As described above, the audio signal processing method according to the present embodiment is an audio signal processing method executed by an audio signal processing device, and includes an obtaining step, a third calculating step, a selection processing step, and a reproducing step.
In the obtaining step, an audio signal is obtained. In the third calculating step, a predetermined sound volume is calculated. The predetermined sound volume is based on a sound volume of a sound at a time at which the sound arrives at a listening position that is a position at which a listener is present. The sound is a sound indicated by the audio signal obtained. The calculating is based on: a predetermined correction characteristic obtained by correcting, using a frequency characteristic indicating auditory sensitivity, a gain characteristic of each predetermined frequency bandwidth related to the sound; and the audio signal obtained. In the selection processing step, whether to output an output signal that is based on the audio signal obtained is selected, based on the predetermined sound volume calculated. In the reproducing step, the output signal is output, when outputting the output signal is selected.
The predetermined sound volume is thus calculated while taking the frequency characteristic indicating the auditory sensitivity into consideration, and whether to output the output signal that is based on the audio signal representing the sound is selected, based on the predetermined sound volume calculated. That is, whether to output the output signal that is based on the audio signal is appropriately selected while taking the auditory sensitivity into consideration. When the output signal is not output, the amount of computation and the computational load are reduced. In other words, it is possible to realize an audio signal processing method that can appropriately reduce the amount of computation and the computational load while taking auditory sensitivity into consideration.
(Supplement)
Note that the aspects understood based on the present disclosure are not limited to the embodiment, and various changes may be performed.
For example, a process performed by a certain constituent element in the embodiment may be performed by another constituent element instead of the specific constituent element. Furthermore, the order of a plurality of processes may be changed, or a plurality of processes may be performed in parallel.
Moreover, ordinals such as first and second used for description may be interchanged, removed, or newly assigned as appropriate. These ordinals do not necessarily correspond to meaningful orders, and may be used to distinguish between elements.
Furthermore, for example, in comparisons between threshold values, “greater than or equal to” a threshold value and “greater than” a threshold value may be read interchangeably. Similarly, “less than or equal to” a threshold value and “less than” a threshold value may be read interchangeably. Moreover, for example, there may be cases in which the terms “time period” and “time” are read interchangeably.
Furthermore, in a process for selecting one or more sounds to be processed from a plurality of sounds, no sounds need be selected as a sound to be processed if no sounds that satisfy the conditions exist. In other words, a case in which no sounds to be processed are selected may be included in the process for selecting one or more sounds to be processed from a plurality of sounds.
Furthermore, at least one of a first element, a second element, or a third element can correspond to the first element, the second element, the third element, or any combination of these.
It should be noted that the information, described in Embodiment 2, indicating the gain characteristic of each predetermined frequency bandwidth related to the reflected sound (indirect sound), the gain characteristic of each predetermined frequency bandwidth related to the direct sound, and the frequency characteristic indicating the auditory sensitivity may be identified based on the spatial information included in the input information, for example. The metadata includes, for example, information representing the reflectance of structures that may reflect sound in the sound space, such as floors, walls, and ceilings, as well as information representing the reflectance of obstacle objects present in the sound space and the reflectance of their reflecting surfaces. Here, reflectance is defined as the ratio of energy or amplitude between reflected sound and incident sound, and is set for each frequency band of a sound. Furthermore, the reflectance may also be referred to as the reflection coefficient. For example, data such as that shown in FIG. 32 may be set in the metadata as information indicating the reflectance. FIG. 32 exemplifies the reflection coefficients for each of the three types of wall surfaces. For a non-sound-emitting object, one reflection coefficient may be set, or a plurality of reflection coefficients may be set. The gain of the equalizer applied to the reflected sound for each frequency band (i.e., the first correction characteristic) may be calculated based on the reflection coefficient set for the reflecting surface related to that reflected sound, described in the above example.
Furthermore, the gain characteristic of each predetermined frequency bandwidth related to a reflected sound (indirect sound), the gain characteristic of each predetermined frequency bandwidth related to a direct sound, and the frequency characteristic indicating the auditory sensitivity may be supplied using a communication line or the like, or may be stored in advance in the memory of analyzer 2301.
In addition, for example, in the embodiment, the case in which the aspects that are understood based on the present disclosure are implemented as an audio signal processing device, an encoding device, or a decoding device has been described. However, the aspects that are understood based on the present disclosure are not limited thereto, and may be implemented as software for executing the audio signal processing method, the encoding method, or the decoding method.
For example, a program for executing the above-described audio signal processing method, encoding method, or decoding method may be stored beforehand in ROM. Then, a CPU may operate according to this program.
Furthermore, a program for executing the above-described audio signal processing method, encoding method, or decoding method may be stored on a computer-readable recording medium. Then, a computer may record, in computer RAM, the program stored on the recording medium, and operate according to this program.
Moreover, each of the above-described constituent elements may be expressed typically as a large-scale integration (LSI), which is an integrated circuit (IC) having an input terminal and an output terminal. These may take the form of individual chips, or all or one or more constituent elements of the embodiment may be encapsulated in a single chip. Depending upon the level of integration, the LSI may be expressed as an IC, a system LSI, a super LSI, or an ultra LSI.
Furthermore, such IC is not limited to an LSI, and a dedicated circuit or a general-purpose processor may be used. Alternatively, a field programmable gate array (FPGA) that allows for programming after the manufacture of an LSI, or a reconfigurable processor that allows for reconfiguration of the connection and the setting of circuit cells inside an LSI may be employed. Furthermore, when a circuit integration technology that replaces LSIs comes along owing to advances in semiconductor technology or to a separate derivative technology, the constituent elements should naturally be integrated using that technology. The adaptation of biotechnology, and the like are also conceivable as possibilities.
Moreover, an FPGA, a CPU, or the like may, by means of wireless communication or wired communication, download all or a part of the software for executing the audio signal processing method, the encoding method, or the decoding method described in the present disclosure. Furthermore, all or a part of software for updating may be downloaded by means of wireless communication or wired communication. Moreover, an FPGA, a CPU, or the like may execute the digital signal processing described in the present disclosure by storing the downloaded software in memory and operating based on the stored software.
At this time, the machine that includes the FPGA, the CPU, or the like may be connected wirelessly or in a wired manner to a signal processing device, or may be connected to a signal processing server over a network. Accordingly, this machine and the signal processing device or the signal processing server may perform the audio signal processing method, the encoding method, or the decoding method described in the present disclosure.
For example, the audio signal processing device, the encoding device, or the decoding device in the present disclosure may include an FPGA, a CPU, or the like. Furthermore, the audio signal processing device, the encoding device, or the decoding device may include: an interface for acquiring, from an external source, the software for causing the FPGA, the CPU, or the like to operate; and memory for storing the acquired software. The FPGA, the CPU, or the like may perform the signal processing described in the present disclosure by operating based on the stored software.
A server may provide the software related to the acoustic processing, the encoding processing, or the decoding processing of the present disclosure. Furthermore, a terminal or a machine may operate as the audio signal processing device, the encoding device, or the decoding device described in the present disclosure by installing the software. Note that the terminal or the machine may install the software by connecting to a server over a network.
Furthermore, the software may be installed on the terminal or the machine by means of another device that is different from the terminal or the machine obtaining data for installing the software by connecting to a server over a network and providing the data for installing the software to the terminal or the machine. Note that VR software or AR software for causing a terminal or a machine to execute the audio signal processing method described by way of the embodiment may be an example of the software.
Note that in the foregoing embodiment, each constituent element may be configured from dedicated hardware, or may be implemented by executing a software program suitable for each constituent element. Each constituent element may be implemented by means of a program executor such as a CPU or a processor loading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
Thus, the device and the like according to one or more aspects have been described by way of the embodiment, but the aspects understood based on the present disclosure are not limited to the embodiment. The one or more aspects may thus include forms obtained by making various modifications to the above embodiments that can be conceived by those skilled in the art, as well as forms obtained by combining constituent elements in different variations, without materially departing from the spirit of the present disclosure.
INDUSTRIAL APPLICABILITY
The present disclosure includes aspects that can be applied to, for example, an audio signal processing device, an encoding device, a decoding device, or a terminal or equipment that includes any of these.
