audio – Devstyler.io https://devstyler.io News for developers from tech to lifestyle Thu, 07 Dec 2023 12:00:10 +0000 en-US hourly 1 https://wordpress.org/?v=6.8.5 Google Introduces Gemini – A New Multimodal AI Model https://devstyler.io/blog/2023/12/07/google-introduces-gemini-a-new-multimodal-ai-model/ Thu, 07 Dec 2023 12:00:10 +0000 https://devstyler.io/?p=115556 ...]]> Google has announced its latest AI model, Gemini, which was designed from the ground up to be multimodal, so it can interpret information in a variety of formats – text, code, audio, image and video.

According to the company, the typical approach to creating a multimodal model involves training components for different formats of information separately and then combining them together. What sets Gemini apart is that it is trained for different formats from the start and then refined with additional multimodal data.

“This helps Gemini seamlessly understand and reason about all kinds of inputs from the ground up, far better than existing multimodal models — and its capabilities are state of the art in nearly every domain,” Sundar Pichai, CEO of Google and Alphabet, and Demis Hassabis, CEO and co-founder of Google DeepMind, wrote in a blog post.

Google also explained that the new model has quite sophisticated reasoning capabilities that allow it to understand complex written and visual information, making it “adept at discovering knowledge that can be difficult to discern among vast amounts of data.”

For example, he can read hundreds of thousands of documents and extract this information to lead to new discoveries in certain fields.

Its multimodal nature makes it particularly suited to understanding and answering questions in complex fields such as mathematics and physics.

Gemini 1.0 offers three distinct versions to cater to various size preferences: Ultra, Pro, and Nano, listed in descending order of size.

According to Google’s initial benchmarking of Gemini, the Ultra version has demonstrated superior performance, outperforming 30 out of the 32 widely used academic benchmarks in model development and research. Notably, Gemini Ultra has achieved a milestone by surpassing human expert performance in massive multitask language understanding (MMLU), encompassing 57 subjects such as math, physics, history, law, medicine, and ethics.

The integration of Gemini Pro into Bard marks a significant milestone, constituting the most substantial update to Bard since its initial launch. The Pixel 8 Pro now leverages Gemini Nano to enhance functionalities like Summarize in the Recorder app and Smart Reply in Google’s keyboard.

Over the coming months, Gemini is set to extend its presence to additional Google products, including Search, Ads, Chrome, and Duet AI.

Starting from December 13, developers can access Gemini Pro through the Gemini API available in Google AI Studio or Google Cloud Vortex AI.

]]>
Uber Introduces Measures Against Unfair Treatment of Drivers https://devstyler.io/blog/2023/11/14/uber-introduces-measures-against-unfair-treatment-of-drivers/ Tue, 14 Nov 2023 09:32:44 +0000 https://devstyler.io/?p=113933 ...]]> Uber is introducing new features aimed at addressing the unfair deactivations often experienced by drivers who transport passengers as well as providers.

The technology that Uber is introducing across the US will identify riders or Uber Eats customers who consistently give bad ratings or reviews in order to get their money back. The company’s blog post says that the claims of these customers will not be taken into account when deciding to rate or disable driver accounts.

The company is also expanding its in-app review center to give drivers and couriers more information about the reasons for deactivating their account, allow them to request a review of the decision, and share additional information, such as audio or video footage.

Last year, Uber introduced a nationwide audio recording feature for drivers and riders. The company also began piloting video recording and said Monday it will expand the pilot to iOS drivers in a dozen U.S. cities, including Atlanta, Denver, Dallas, Minneapolis and select drivers in Los Angeles.

For an extended period, drivers in the app-based gig economy have been expressing their opposition to unjust deactivations, essentially tantamount to termination. Numerous drivers have participated in collective legal actions against the company. Their grievances include allegations that certain riders file malicious or biased complaints. Additionally, drivers contend that they face a lack of transparency in accessing the details of these complaints, hindering their ability to challenge them. Furthermore, Uber is criticized for providing minimal avenues for drivers to dispute these claims.

Verifying riders and other updates for drivers
Uber also said that in 2025, the company will expand verification of rider identities. Riders will be identified based on simple third-party checks — like if your name matches the credit card on file. If Uber can’t verify a rider’s identity that way, they’ll ask for an ID, but that won’t be the standard. Uber wouldn’t share further information on this, such as whether Uber will automatically verify riders or whether riders will have to opt in.

Uber is also integrating Android Auto with the Uber Driver app, allowing drivers who use Androids to now see heat maps, accept trips and use on-screen navigation from their car dashboard. The integration comes several months after Uber launched something similar with Apple CarPlay in February.

Finally, Uber added a tool in the app to help couriers find nearby parking. The company said it will also add map labels that specify exact drop-off doors or photos of the building to make it clearer to couriers where a customer requested food to be dropped off.

]]>
Does it violate the Law? Tech Giants Obliterated Protest Song https://devstyler.io/blog/2023/06/15/does-it-violate-the-law-tech-giants-obliterated-protest-song/ Thu, 15 Jun 2023 07:51:33 +0000 https://devstyler.io/?p=107794 ...]]> A popular protest song from Hong Kong is no longer available on several music streaming platforms, including Apple Music and Spotify, after the city government issued a court order banning the tune, Nikkei Asia reports.

The song Glory to Hong Kong can’t even be found in Meta’s Instagram audio pictures. For now, the song is only available on YouTube.

Is the performance a mistake?
Authorities in Hong Kong are trying to ban the pro-democracy song after organisers of several international sporting events “mistakenly” performed it instead of China’s national anthem. However, the removal of the song from music platforms comes about a month before the Supreme Court is due to rule on 21 July – a possible example of pre-emptive self-censorship that could set a new precedent.

The reason?
The band DGX Music, which is also the creator of the song, announced that it had encountered technical issues unrelated to the streaming platforms and apologized for the temporary service disruption.

Clues to the crime?
“Glory to Hong Kong” was the unofficial anthem of protesters during the mass demonstrations in 2019. However, after a court found that the lyrics “Free Hong Kong, revolution of our time” could incite a crime under Beijing’s imposed national security law the song became illegal.

Are human rights being violated?
Activist groups said banning the song violates international human rights law and further undermines freedom of expression in the former British colony, which has already blocked access to several websites deemed a threat to national security. The reduction of space for free expression is seen as undermining Hong Kong’s reputation as an international business hub.

“If big platforms like Google decide to leave Hong Kong due to regulatory concerns, it will certainly give Hong Kong a big hit in terms of global investors’ confidence,”

said George Chen, managing director for The Asia Group, a Washington-headquartered business and policy consulting firm.

Charles Mok, a visiting scholar at Stanford University’s Cyber Policy Center and a former legislator in Hong Kong, said tech platforms have to take into account geopolitical tensions between the U.S. and China. Avoiding potential repercussions from U.S. politicians would be one major factor, he added.

In mainland China, the government maintains complete control over the internet and censorship is widespread.

]]>
Four ML Trends to Adapt to in the Future https://devstyler.io/blog/2023/03/14/four-ml-trends-to-adapt-to-in-the-future/ Tue, 14 Mar 2023 07:42:51 +0000 https://devstyler.io/?p=102999 ...]]> Over the next six years, the global ML market size is expected to grow from $21.17 billion in 2022 to $209.91 billion in 2029. Expecting this growth means that in 2023, organizations will see a paradigm shift in how they prioritize ML investments, Spiceworks reports.

Most companies say they are using six different tools to build models and learn, with executives focusing more on downstream ML capabilities such as observability and function management. The shift from building complex, company-wide ML models to smaller, task-focused models increases their portability and reduces barriers to market. And today we’ve chosen to introduce you to the top four ML trends that Spiceworks presents.

Generative AI
Generative AI can create new content, including audio, code, images, text, simulations, and videos. It uses deep neural networks with billions of parameters to enable complex pattern recognition.

Unsupervised or semi-supervised learning algorithms are making huge strides in accelerating research and development (R&D) cycles in the medical and financial forecasting fields.

The sector will be increasingly regulated. The EU’s AI law, the US Privacy and Data Protection Act and the Open Source Software Assurance Act are all breaking the mould to promote the safety and security of modern technological lifestyles. Whether your company has entered the world of ML or not, enterprises should be aware of these acts and plan strategies to strengthen fraud detection to mitigate risks against the latest ML tools.

Computer vision
Computer vision (CV) accounts for the largest share of the AI and ML market. It is a field of AI that can capture, process and analyze real-world images, enabling the extraction of meaningful, contextual information.

One of the sectors where CV is making an impact is the automotive industry. It can detect defects in the bodywork of cars and underpin the development of applications such as self-driving cars. High-resolution cameras with background CV systems identify surrounding objects, people and movements that automatically trigger the vehicle’s response.

Increasingly, it will help maintenance service providers perform efficient and thorough inspections by using cameras to identify dents and mechanical parts that are out of place. Using CV, engineers can process the images and identify discrepancies within seconds.

Cross-industry AI Synergy
ML will be less concerned with company-specific models and more with data-driven models that can be used across sectors.

Doctors and scientists are experimenting with ML and CV technologies, training them to recognize and classify rare genetic skin conditions. Specialists walk the aisles, in some cases hourly, to count merchandise and ensure product availability. But what if they put CV apps on the shelves to help them monitor real-time inventory?

As companies begin to share the investment cost of tools that can analyze visual patterns and detect everything from rare diseases to product movement, more experimentation and affordable models can be created.

Late adopters are increasingly looking at use cases from more mature ML industries, such as automotive and healthcare, and adapting them to support their business needs.

ML Data Scientist Upskilling vs. Low-code solutions

Codeless and low-code (LCNC) platforms allow users with or without programming language knowledge to manage and build ML tools through intuitive interfaces such as point-and-click and drop-down menus.

ML use cases are expanding rapidly as developers begin to reinvent workflows based on what the technology can provide. AI people can re-engineer systems by pre-programming them to proactively deliver alerts and red flags to users when certain triggers are reached.

Nonetheless, LCNC tools are fundamentally limited in the scope of customization by design. Highly skilled software engineers will be needed to build, monitor, and scale these platforms. The latter needs are likely to lead to a demand for distinct new positions, such as human-computer interaction managers.

]]>
Meta Quest Improves Spatial Audio https://devstyler.io/blog/2023/02/13/meta-quest-improves-spatial-audio/ Mon, 13 Feb 2023 10:11:19 +0000 https://devstyler.io/?p=101101 ...]]> Meta has added immersive audio capabilities to its proprietary Presence Platform. The “XR Audio SDK” is designed to make it easier for developers to incorporate spatial, localized audio, Developer Oculus reports.

It’s currently only available for the Unity engine, which is widely used in VR. Support is planned for Unreal Engine, Wwise, and FMOD.

Presence Platform
The Presence Platform is a collection of development tools and programming interfaces that enable hands and voice interaction and augmented reality capabilities with Quest 2 and Quest Pro, for example.

Applications of the new immersive audio features include virtual reality, augmented reality and mixed reality. For the latter, Meta’s Quest 2 and Quest Pro VR headsets add computer graphics to the video image from their front-facing cameras.

In addition to Meta devices, the new audio SDK supports “almost any standalone mobile VR device” as well as PC VR (e.g. Steam VR) and third-party devices.

New features
New features include better handling of the head, outer ear, and torso filtering effects that greatly affect sound in the real world: The Head-Related Transfer Function (HRTF) is designed to mimic authentic audio perception accurately. Without it, sounds in your immediate environment will sound unnatural.

The Spatial Audio Rendering and Room Acoustics features build on the previous Oculus Spatializer and will continue to be developed. The system is much better suited for use in VR than the built-in audio systems in popular game engines, which are primarily designed for consoles and PCs, Meta said in its developer blog.

The new Audio SDK offers both flexibility and ease of use. According to Meta, even developers with no audio experience will be able to mix audio, which is essential for immersion.

What else we need to know
The previous Oculus Spatializer will continue to be supported in Unreal Engine, FMOD or Wwise or for those who prefer a native API solution. Meta does not recommend upgrading in these areas yet.

For new projects in Unity Engine, Meta recommends using the new “XR Audio SDK” to better maintain applications in the long run or to try out experimental features.

]]>
Sonos CEO Criticizes Amazon and Google https://devstyler.io/blog/2023/02/10/sonos-ceo-criticizes-amazon-and-google/ Fri, 10 Feb 2023 13:26:22 +0000 https://devstyler.io/?p=100907 ...]]> During Sonos’ earnings call for the first quarter of 2023, CEO Patrick Spence took a clear swipe at competitors in the big tech sector, including Amazon, Google and Apple, for lacking creativity in recent months, The Verge reports.

None of Sonos’ main rivals in smart speakers released new audio hardware during the holiday quarter. And Apple’s second-generation HomePod appeared after it had already finished. But the company doesn’t seem as impressed with the devices Apple is offering.

Discounts on Sonos products during the holidays helped the company deliver strong earnings and beat Wall Street expectations for its fiscal first quarter. This is the first time Sonos has been able to offer these promotions in three years.

“We’ve gone through fiscal Q1, which is the height of the consumer electronics and audio season and, you know, it was… we’ve seen some of the traditional players go heavy discounting and, kind of like a traditional playbook for C.E. that, you know, we’ve always fought against and don’t really believe in. And then, you know, the big tech players, we just haven’t seen them active and we haven’t seen them, you know, doing anything interesting.”

Said Patrick Spence.

Google’s most recent home audio product is the aging Nest Audio, and Amazon’s fall 2022 hardware event went light on smart speakers: the company introduced a refreshed Echo Dot and added a new white color option for the top-of-the-line Echo Studio. The regular Echo speaker was last refreshed in 2020.

“I just feel like there’s others that are probably questioning their investments in this area, and we are investing in four new categories. We are going to raise the bar in our existing categories. I mean, we’ve got a lot going on,”

Spence added.

Sonos has said it will release a product in the first of those new categories sometime in 2023.

The company has placed greater emphasis on bundles that increase the percentage of households using multiple products. The average number of products for each Sonos household is currently 2.98. Sonos sees a $5 billion revenue opportunity if it can successfully transition single-product customers to a multi-device lifestyle.

]]>
Stability AI Doubles in AWS https://devstyler.io/blog/2022/12/01/stability-ai-doubles-in-aws/ Thu, 01 Dec 2022 11:47:48 +0000 https://devstyler.io/?p=95163 ...]]> Stability AI is doubling down on its cloud, making the cloud service provider of choice for building and scaling its AI models to generate images, languages, audio, video and 3D content, AWS announced.

Stability AI will work with AWS to make its open source tools and model available to more students, researchers, startups and enterprises.

“At Stability AI, our mission is to build the foundation to activate humanity’s potential through AI. AWS has played an integral role in scaling our open-source foundation models across modalities. We are delighted to run these models on Amazon SageMaker to enable thousands of developers and millions of users to leverage the power of AI with a robust set of tools”

 

said Emad Mostaque, founder and CEO of Stability AI.

 

“We look forward to seeing the amazing things that developers build and customers design and implement using collective intelligence and augmented technology”

he continued.

Stability AI plan to use AWS’ SageMaker ML platform, on top of its lower-level infrastructure services with GPUs and AWS’ own Trainium chips. Stability AI’s open source approach seems to be winning in bringing on more developers and driving innovation in this space.

 

]]>
Data2vec – The Multi-Modal AI Algorithm  https://devstyler.io/blog/2022/02/23/data2vec-the-multi-modal-ai-algorithm/ Wed, 23 Feb 2022 08:09:48 +0000 https://devstyler.io/?p=81587 ...]]> Meta AI recently open-sourced data2vec, a unified framework for self-supervised deep learning on images, text, and speech audio data. When evaluated on common benchmarks, models trained using data2vec perform as well as or better than state-of-the-art models trained with modality-specific objectives, noted InfoQ.

Data2vec is a framework that uses the same learning method for either speech, NLP or computer vision. The core idea is to predict latent representations of the full input data based on a masked view of the input in a self-distillation setup using a standard Transformer architecture, told arXiv.

According to a post in Meta’s blog, data2vec is simplifying the different algorithms functioning by training models to predict their own representations of the input data, regardless of the modality. A single algorithm can work with completely different types of input. This removes the dependence on modality-specific targets in the learning task. Directly predicting representations is not straightforward, and it requires defining a robust normalization of the features for the task that would be reliable in different modalities.

]]>
Google Announced Nest Hub 2nd Gen In India https://devstyler.io/blog/2022/01/12/google-announced-nest-hub-2nd-gen-in-india/ Wed, 12 Jan 2022 14:15:33 +0000 https://devstyler.io/?p=78721 ...]]> On Wednesday Google announced that its Nest Hub (2nd Gen) is now available in the country at Rs 7,999.

The Google Nest Hub (2nd Gen) is available in chalk and charcoal colours across Flipkart, Tata Cliq and Reliance Digital. In a statement the company said:

“The new Nest Hub’s speaker is based on the same audio technology as Nest Audio and has 50 per cent more bass than the original Nest Hub.”

The Nest Hub (2nd Gen) features an edgeless glass display. It is designed with recycled materials with its plastic mechanical parts containing 54 per cent recycled post-consumer plastic. The company added:

“And with an additional mic, it hears you better than ever, resulting in a more responsive Google Assistant.”

By default, the user’s audio recordings are not retained. All of the recent activity can be deleted. The only necessity here is to say something like “Ok Google, delete everything I said last week”.

Google said that the new Nest Hub can fill any room with music from services like YouTube Music, Spotify, Apple Music, Gaana and JioSaavn. It can also play movies, videos and TV shows with a subscription from providers like Netflix, and YouTube Premium.

]]>
AI Chatbots Can Now Tell How You’re Feeling https://devstyler.io/blog/2022/01/07/ai-chatbots-can-now-tell-how-you-re-feeling/ Fri, 07 Jan 2022 09:52:38 +0000 https://devstyler.io/?p=78380 ...]]> The new discovery uses AI to break down sentences into sounds and tones. According to Five9 their new technology will bring big savings to companies on personnel costs.

The key here is using a human voice to train the AI. Callan Scheballa, a project manager with Five9, explained:

“At the end of the day people are going to understand that they are talking to a machine, they are going to understand that they are talking to software.”

He added:

“Yet, the voice that it uses, there’s no reason for it to not sound great. Now, the more. lifelike that it can sound, at least in our experience, the better the reception of the customers that are going to be talking to it.”

Five9 auditioned actors in London and decided on Joseph Vaughn to record a series of scripts for the company, in turn to find their latest voice.

That audio was then broken down into sounds and tones rather than words. It’s enabling the AI program to recreate not only sentences but also distinct moods.

The software is trained to detect word combinations and tones in the caller’s voice so that it can respond to their emotional state as well. Rhyan Johnson, an engineer from Wellsaid labs who is involved in the project, said:

“We are capturing all of the audio data and all of the combination of frequencies and vibrations that are inherent to a voice, that as a human we would recognise is a voice, but the machine just is guessing sounds.”

The company says their AI agents have already answered more than 82 million calls for healthcare providers, large retailers as well as insurance companies, banks, local businesses, and state and local governments.

The new Virtual Voiceover tech will be available next year.

]]>