Abstract:
A fast and efficient acoustic feature super vector generation method was proposed to effectively improve the recognition accuracy and speed yielded by traditional frame based acoustic features. This paper makes 3 contributions:firstly, certain number of acoustic feature vectors extracted from continuous audio frames was combined to be an acoustic feature image; secondly, AdaBoost. MH algorithm was used to select higher representative 2D-Haar pattern combinations to construct super feature vectors; thirdly, random feature selection method was proposed to further improve the processing speed. Experimental results show that under 3 kinds of audio recognition occasions such as audio events recognition, speaker recognition, speaker gender recognition, the use of 2D-Haar acoustic feature super vector can make SVM, C5.0, AdaBoost algorithms obtain higher recognition accuracy than ones that MFCC, PLP, LPCC and other traditional acoustic features yielded, and can make the training processing 7~20 times faster and the recognition processing 5~10 times faster.