Enhancement of throat microphone recordings using gaussian mixture model probabilistic estimator

Turan, Mehmet Ali Tuğtekin

dc.contributor.advisor	Erzin, Engin
dc.contributor.author	Turan, Mehmet Ali Tuğtekin
dc.date.accessioned	2020-12-08T07:49:48Z
dc.date.available	2020-12-08T07:49:48Z
dc.date.submitted	2013
dc.date.issued	2018-08-06
dc.identifier.uri	https://acikbilim.yok.gov.tr/handle/20.500.12812/168843
dc.description.abstract	Gırtlak mikrofonu, ses tellerindeki titreşimi gırtlaktan gelen sinyallerle beraber ileten ve kullanan kişinin boynuna taktığı insan bedeniyle temas eden bir mikrofon türüdür. Bu bağlantı sayesinde, titreşimleri havadan alan akustik mikrofonlara nazaran gürültü gibi çevresel etmenlere karşı daha gürbüz bir iletişim sağlar. Gırtlak mikrofonu ile kaydedilen sesler kısmen de olsa anlaşılmasına rağmen, doğal olmayan ve kulağı rahatsız edici bir yapıdadır. İşte bu çalışma gırtlak mikrofonlarındaki üretilemeyen frekans aralıklarını geri kazanabilmeyi amaçlarken aynı zamanda sesin kaynak ve süzgeç kısımlarını doğru tahmin edebilme sorununu, gırtlak ve akustik kayıtları müşterek bir şekilde çözümleyerek irdelemektedir. Bu bağlamda, ortalama kare hatasınıen aza indirerek, ses birimlerine bağlı Gauss karışım modeli tabanlı bir kestirici sistemi öne sürülmüştür. Kaynak-süzgeç ayrıştırması çerçevesinde, gırtlak ve akustik süzgecinin görüngesel farklılıklarının, gırtlak mikrofonundan gelen ses kalitesini düşüren önemli bir etmen olduğunu gözlemledik. Bu sebepten ötürü, yukarıda bahsedilen farkı görüngesel eğim vektörü olarak modelleyip, gırtlak süzgecini iyileştirici bir sistemi ayrıca öne sürdük. Ortaya konulan sistemlerin katkılarını yorumlayabilmek için hem nesnel hem de öznel deneyler tasarladık. Nesnel deneyler, logaritmik görünge tahribatı ve ses kalitesinin algısal değerlendirilmesi kıstasları üzerinden incelendiler. Bununla birlikte, öznel değerlendirmeler ise A/B eş karşılaştırma deneyi şeklinde tatbik edildi. Hem nesnel hem de öznel deneyler gösterdi ki öne sürülen ses birimi tabanlı kestirimler, halihazırda bulunan Gauss karışım modeli tabanlı kestirimlere göre tutarlı bir şekilde iyileştirmeler sağlamaktadır.
dc.description.abstract	The throat microphone is a body-attached transducer that is worn against the neck. It captures the signals that are transmitted through the vocal folds, along with the buzz tone of the larynx. Due to its skin contact, it is more robust to the environmental noise compared to the acoustic microphone that picks up the vibrations through air pressure, and hence the all interventions. The throat speech is partly intelligible, but gives unnatural and croaky sound. This thesis tries to recover missing frequency bands of the throat speech and investigates envelope and excitation mapping problem with joint analysis of throat- and acoustic-microphone recordings. A new phone-dependent GMM-based spectral envelope mapping scheme, which performs the minimum mean square error (MMSE) estimation of the acoustic-microphone spectral envelope, has been proposed. In the source-filter decomposition framework, we observed that the spectral envelope difference of the excitation signals of throat- and acoustic-microphone recordings is an important source of the degradation in the throat-microphone voice quality. Thus, we also model spectral envelope difference of the excitation signals as a spectral tilt vector, and propose a new phone-dependent GMM-based spectral tilt mapping scheme to enhance throat excitation signal. Experimental evaluations are performed to compare the proposed mapping scheme using both objective and subjective evaluations. Objective evaluations are performed with the log-spectral distortion (LSD) and the wide-band perceptual evaluation of speech quality (PESQ) metrics. Subjective evaluations are performed with A/B pair comparison listening test. Both objective and subjective evaluations yield that the proposed phone-dependent mapping consistently improves performances over the state-of-the-art GMM estimators.	en_US
dc.language	English
dc.language.iso	en
dc.rights	info:eu-repo/semantics/openAccess
dc.rights	Attribution 4.0 United States	tr_TR
dc.rights.uri	https://creativecommons.org/licenses/by/4.0/
dc.subject	Elektrik ve Elektronik Mühendisliği	tr_TR
dc.subject	Electrical and Electronics Engineering	en_US
dc.title	Enhancement of throat microphone recordings using gaussian mixture model probabilistic estimator
dc.title.alternative	Gırtlak mikrofonu kayıtlarının gauss karışım modeli aracılığıyla iyileştirilmesi
dc.type	masterThesis
dc.date.updated	2018-08-06
dc.contributor.department	Elektrik-Elektronik Mühendisliği Anabilim Dalı
dc.subject.ytm	Gaussian mixture model
dc.identifier.yokid	10014915
dc.publisher.institute	Fen Bilimleri Enstitüsü
dc.publisher.university	KOÇ ÜNİVERSİTESİ
dc.identifier.thesisid	332230
dc.description.pages	53
dc.publisher.discipline	Diğer

Files in this item

Name:: yokAcikBilim_10014915.pdf
Size:: 1.925Mb
Format:: PDF
Description:: File_10014915

View/Open

This item appears in the following Collection(s)

TEZLER

Show simple item record

Except where otherwise noted, this item's license is described as info:eu-repo/semantics/openAccess