Showing posts with label speech. Show all posts
Showing posts with label speech. Show all posts

Tuesday, 16 June 2009

CFP: 3rd Workshop on Learning the Semantics of Audio Signals (Graz, Austria)

3RD WORKSHOP ON LEARNING THE SEMANTICS OF AUDIO SIGNALS (LSAS) 2009
SAMT 2009 CONFERENCE, GRAZ, AUSTRIA
4 DECEMBER 2009
CALL FOR PAPERS

The Workshop on Learning the Semantics of Audio Signals (LSAS) focuses especially on researchers that are working on semantics, description, representation and understanding of music with the goal to intensify the exchange of ideas between the different research communities involved, to provide an overview of current activities in this area and to point out connections between them. Therefore, the topics of interest include, but are not limited to the following aspects with respect to music
information retrieval systems:

- Audio signal processing and feature extraction
- Content-based audio retrieval
- Music perception, cognition, affect and emotions
- Music structure analysis
- Semantic audio description and analysis
- Ontologies for music and sound description
- Standards for audio content description
- Machine learning methods for feature extraction and mapping
- Personalization of music retrieval systems

While our main focus is on music data, we are also interested on work related to other audio data such as speech.

Important dates:
- Deadline for paper submission: September 7, 2009
- Notification of acceptance: September 21, 2009
- Camera-ready paper submission: October 12, 2009

For the full Call For Papers and further information on the workshop, please visit http://lsas2009.dke-research.de

Friday, 12 June 2009

PhD Studentship Audio-Visual Machine Listening (Queen Mary, London)

PHD STUDENTSHIP IN AUDIO-VISUAL MACHINE LISTENING
CENTRE FOR DIGITAL MUSIC, QUEEN MARY UNIVERSITY OF LONDON
DEADLINE: 1 JULY 2009

Applications are invited for a 3-year PhD studentship to undertake research into audio-visual analysis of sound events, with an emphasis on non-speech (music and environmental) sounds.

Research in audio and video signal processing has traditionally taken place in different research groups, but audio-visual processing is now increasingly being linked. In speech processing, for example, image processing methods for lip reading can help improve speech recognition performance, particularly in noisy environments. One promising technique for audio-visual analysis of particular interest in this project is "sparse representations", which looks for a representation of an audio or video signal using a small number of non-zero elements, which can then be associated together.

The work will form part of a programme of research in "Machine Listening using Sparse Representations", supported by a Leadership Fellowship from the UK Engineering and Physical Sciences Research Council (EPSRC). This is a concerted programme of research aiming to establish machine listening as a key enabling technology to improve our ability to interact with the world.

The successful candidate will be based in the Machine Listening lab of the world-leading Centre for Digital Music at Queen Mary University of London, working under the supervision of Prof Mark Plumbley (www.elec.qmul.ac.uk/people/markp).

Candidates should have a first or upper second honours degree or equivalent in electronic engineering, mathematical science, physics, statistics, computer science, or allied disciplines, and be able to demonstrate excellent mathematical and programming skills. Research experience in digital signal processing of audio, image and/or video signals is desirable.

The studentship is for 3 years (subject to satisfactory progress) starting from 1 September 2009 or as soon as possible thereafter. It includes tuition fees and a tax-free stipend in line with EPSRC recommendations (currently at £14,940 per annum for the 2008/09 session), and is open to EU (including UK) candidates.

For informal enquiries, please contact Prof Mark Plumbley, Queen Mary University of London, Email: mark.plumbley [at]elec.qmul.ac.uk.

How to apply

To apply for the PhD studentship please email the following documents to Prof Mark Plumbley: mark.plumbley@elec.qmul.ac.uk: a completed application form, a CV listing all publications, your representative publications in PDF format, two independent reference letters, and other relevant documents as requested (see http://www.eecs.qmul.ac.uk/phd/apply.php). These documents must also be provided in paper form and sent to the Admissions and Recruitment Office (see the application form for address).

The closing date for applications is Wednesday 1st July 2009.

Friday, 1 May 2009

Job Posting: Research and Development Position in Speech Recognition, Processing and Synthesis (IRCAM)

JOB POSTING: RESEARCH AND DEVELOPMENT POSITION IN SPEECH RECOGNITION, PROCESSING AND SYNTHESIS
IRCAM, PARIS

The position is available immediately in the Speech group of the Analysis/Synthesis team at Ircam. The Analysis/Synthesis team undertakes research and development centered on new and advanced algorithms for analysis, synthesis and transformation of audio signals, and, in particular, speech.

JOB DESCRIPTION:
A full-time position is open for research and development of advanced statistics and signal processing algorithms in the field of speech recognition, transformation and synthesis.
http://www.ircam.fr/anasyn.html (projects Rhapsodie, Respoken, Affective Avatars, Vivos, among others)

The applications in view are, for example,
- Transformation of the identity, type and nature of a voice
- Text-to-Speech and expressive Speech Synthesis
- Synthesis from actor and character recordings.

The principal task is the design and the development of new algorithms for some of the subjects above and in collaboration with the other members of the Speech group. The research environment is Linux, Matlab and various scripting languages like Perl. The development environment is C/C++, for Windows in particular.

REQUIRED EXPERIENCE AND COMPETENCE:
O Excellent experience of research in statistics, speech and signal processing
O Experience in speech recognition, automatic segmentation (e.g. HTK)
O Experience of C++ development
O Good knowledge of UNIX and Windows environments
O High productivity, methodical work, and excellent programming style.

AVAILABILITY:
The position is available in the Analysis/Synthesis team of the Research and Development department of Ircam to start as soon as possible.

DURATION:
The initial contract is for 1 year, and could be prolonged.

EEC WORKING PAPERS:
In order to be able to begin immediately, the candidate SHALL HAVE valid EEC working papers.

SALARY:
According to formation and experience.

TO APPLY:
Please send your CV describing in a very detailed way the level of knowledge, expertise and experience in the fields mentioned above (and any other relevant information, recommendations in particular) preferably by email to:

Xavier.Rodet@ircam.fr (Xavier Rodet, Head of the Analysis/Synthesis team)
Or by fax: (33 1) 44 78 15 40, attention of Xavier Rodet
Or by post to: Xavier Rodet, IRCAM, 1 Place Stravinsky, 75004 Paris, France

TITRE DU POSTE :
CDD EN RECHERCHE ET DEVELOPPPEMENT EN RECONNAISSANCE, TRAITEMENT ET
SYNTHESE DE LA PAROLE

Le poste est disponible immédiatement dans le groupe Parole de l'équipe Analyse/Synthèse de l'Ircam. L'équipe Analyse/Synthèse mène des recherches et développements centrés sur des algorithmes nouveaux et avancés pour l'analyse, la synthèse et la transformation des signaux sonores, et en particulier de la parole.

DESCRIPTION DU POSTE :
Un poste à plein temps est ouverte en recherche et développement d'algorithmes avancés de statistique et de traitement du signal, pour le reconnaissance, le traitement et la synthèse de la parole. http://www.ircam.fr/anasyn.html (projets Rhapsodie, Respoken, Affective Avatars, Vivos, entre autres)

Les applications visées sont, par exemple,
- Transformation d'identité, de type et de nature de voix
- Synthèse de voix expressive, à partir du texte
- Synthèse à partir de corpus d'acteurs et de personnages.
La tâche principale est la conception et le développement de nouveaux algorithmes pour certains des sujets ci-dessus et en collaboration avec les autres membres du groupe Parole.

L'environnement de recherche et Linux, Matlab et divers langages de scripting comme Perl. L'environnement de développement est C/C++ pour Windows notamment.

EXPÉRIENCE ET COMPÉTENCE REQUISES :

O Excellente expérience de recherche en statistique, traitement de la parole et du signal
O Expérience en reconnaissance, segmentation automatique de la parole (e.g. HTK)
O Expérience de développement en C++
o Bonne connaissance des environnements UNIX et Windows
O Haute productivité, travail méthodique, et excellent style de programmation.

DISPONIBILITÉ :
Le poste est disponible dans l'équipe "Analyse/Synthèse" du département Recherche et Développement de l'Ircam pour commencer aussitôt que possible.

DURÉE :
Le contrat initial est de 1 an, et pourra être prolongé.

AUTORISATION DE TRAVAIL DANS LA CEE :
Afin de pouvoir commencer immédiatement le candidat DOIT AVOIR un permis de travail dans la CEE.

SALAIRE : Selon formation et expérience.

CANDIDATURE :
Priere d'envoyer un CV décrivant de façon très détaillée le niveau de connaissance, d'expertise et d'expérience dans les domaines mentionnés ci-dessus (ainsi que tout autre information pertinente,recommandations en particulier) de préférence par e-mail à :

Xavier.Rodet@ircam.fr (Xavier Rodet, responsable de l'Équipe Analyse/Synthèse)
Ou par fax à : (33 1) 44 78 15 40, à l'attention de Xavier Rodet
Ou par voie postale à :
Xavier Rodet, IRCAM, 1 Place Stravinsky, 75004 Paris, France
 
Creative Commons License
Interesting Music Stuff (IMS) is licensed under a Creative Commons Licence. Any redistribution of content contained herein must be properly attributed with a hyperlink back to the source.
Click on the time link at the bottom of the post for the direct URL
and cite Colin J.P. Homiski, Interesting Music Stuff.