Google Sounds off on New Audio Search Tool

Techwriter10 0 Tallied Votes 263 Views Share

Google quietly introduced an audio search tool called this week in Google Labs. For now, Google is using to the tool for experimentation purposes to index political content on YouTube videos, but chances are they are exploring this for more than good citizenship points and will expand it at some point in the future. Arnaud Sahuguet, the Product Manager for the Google Audio Indexing project explained to me exactly what this new development means.

Up until now, only a few small multimedia search companies have been able to index audio effectively including Nexidia and TVEyes, two pioneering companies in the audio indexing space. In fact I wrote about these companies in a Streaming Media Magazine article last year called . These companies are able to index with incredible speed with Nexidia building a phonetic index and TVEyes using a hybrid approach of phonetic indexing and using a library of known terms.

According to Sahuget, Google Audio Indexing uses speech technology to transform spoken words into text and uses the Google indexing technology to return the best results to the user. "The returned videos are ranked based -- among other things -- on the spoken content, the metadata, and the freshness of the video. We periodically crawl the YouTube political channels for new content. As soon as a new video is uploaded to YouTube, it is processed by our system and made available in our index for users to search."

When a a couple of years ago, I wondered at the time what it intended to do with it, especially after investing $1.6 billion. With so many people looking at YouTube video, it certainly had potential to generate substantial advertising revenue, but beyond monetization of the medium, Google is first and foremost a search company, so developing some way to search video makes a lot of sense. Sahuget says that more and more content is being created in online video formats, which could explain why Google has decided to develop this tool now.

"Speech-to-text transcription is a useful tool which enables users to instantly find and consume video content by searching across videos for specific terms, even if the video's publisher hasn't transcribed it themselves." He explains that it's time-consuming for users to watch an entire video when they are only looking for a certain part. "Using speech-to-text technology, we can identify the portions of the video where the relevant content is spoken."

For now Sahuget says they are limiting the project to the YouTube political content, but it seems unlikely it will stop there even though he would not commit to anything specific. "Speech recognition is a challenging problem and we are constantly working to improve our technology, but we have nothing to announce at this time," he says. I'm betting this is only the beginning.

Dani AI

Generated

Smart follow-up to this thread from and — the ability to make spoken words searchable shifts priorities for both creators and SEOs. Beyond the excitement, there are practical steps and pitfalls that weren’t fully explored in the early posts. The checklist below focuses on things anyone publishing video should do now, and why each item matters.

Provide a usable transcript and clean captions. Auto-generated captions are a starting point; they should be edited for accuracy and uploaded as a separate file so crawlers and users can read the text without playing the video. Include timecodes and simple speaker labels in the transcript so search results can point to the right moment. Put the full transcript on the same page as the video rather than burying it only inside the player.

Improve discoverability and user experience. Put the most important phrases in the title and the first two lines of the description, add chapter timestamps, and host a text summary or bullets above the transcript. Keep audio clear: minimize background music during speech, separate voices where possible, and use short, explicit signposting (“In this section I explain X”) so automated systems pick up key phrases correctly.

Watch for accuracy, privacy and legal issues. Speech recognition makes false positives and mis-transcriptions more likely for niche terms, names, or heavy accents — check analytics and sample queries to see what’s being picked up. Don’t publish verbatim transcripts of third-party private or copyrighted material without permission. For high-value or sensitive content, prefer human transcription or careful review of auto-captions to avoid misleading indexing or legal exposure.

These steps will help bridge the gap between raw audio and useful search results while reducing common pitfalls that early adopters tend to run into.

acejames1 0 Junior Poster

that was an amazing idea and noone had a clue til someone let the cat out of the bag and roam free but now we see the true thoughts behind google. it will soon be a contender for lots of more things but that is amazing to be able to search youtube videos.

Techwriter10 42 Practically a Posting Shark

Absolutely, it's huge and it's been a huge gap on the internet to date. Lots of video, but no efficient *mainstream* way to search inside the video to get to the bits you really need or want to see.

If they pull this off so that they can index video audio tracks as quickly and efficiently as they do text, it will enable us as searchers to search across media without regard to whether it's text or video and that will be a major leap forward. But still a long way to get there. This is just a first step.

Techwriter10 42 Practically a Posting Shark

Chris Sherman from SearchWise () suggests taking a look at EveryZing (http://www.everyzing.com/) for a company doing some very interesting work in the video search space.

Be a part of the DaniWeb community

We're a friendly, industry-focused community of developers, IT pros, digital marketers, and technology enthusiasts meeting, networking, learning, and sharing knowledge.