back
259 comments
I've tested bard/gemini extensively on tasks that I routinely get very helpful results from GPT-4 with, and bard consistently, even dramatically underperforms.

It pains me to say this but it appears that bard/gemini is extraordinarily overhyped. Oddly it has seemed to get even worse at straightforward coding tasks that GPT-4 manages to grok and complete effortlessly.

The other day I asked bard to do some of these things and it responded with a long checklist of additional spec/reqiurement information it needed from me, when I had already concisely and clearly expressed the problem and addressed most of the items in my initial request.

It was hard to say if it was behaving more like a clerk in a bureaucratic system or an employee that was on strike.

At first I thought the underperformance of bard/gemini was due to Google trying to shoehorn search data into the workflow in some kind of effort to keep search relevant (much like the crippling MS did to GPT-4 in it's bingified version) but now I have doubts that Google is capable of competing with OpenAI.

I don't think Google has released the version of Gemini that is supposed to compete with GPT4 yet. The current version is apparently more on the level of GPT 3.5, so your observations don't surprise me
On the flip side, I find that GPT4 is constantly getting degraded. It intentionally only returns partial answers even when I direct it specifically not to do so. My guess is, that they are trying to save on CPU consumption by generating shorter responses.
My personal favorite Bard failure mode is when I need help with Google Cloud and Bard has no idea what to do but GPT tells me *exactly* what I need.

If you can’t even support your own products…I’m not sure what I’m supposed to do with this pos.

> I've tested bard/gemini extensively on tasks that I routinely get very helpful results from GPT-4 with, and bard consistently, even dramatically underperforms.

Yes. And I don't buy the lmsys leaderboard results where Google somehow shoved a mysterious gemini-pro model to be better than GPT-4. In my experience, its answers looked very much like GPT-4 (even the choice of words) so it could be that Bard was finetuned on GPT-4 data.

Shady business when Google's Bard service is miles behind GPT-4.

I guess Pro is not supposed to be on par with GPT4. That would be Ultra coming out sometime in the first quarter. I’m going to reserve judgement till that is released.
Bard has been dead to me the second I saw it was not available in Canada... GPT all the way to be honest.
How were you able to test Gemini Pro before today? Are you able to test Gemini Ultra?
My experience is the opposite. I'm really tired of fighting ChatGPT.
By comparison I find bing image generator kicks dall-es ass
I ran the obligatory "astronaut riding a horse in space" prompt initially, and was returned two images -- one which was well composed and another which appeared to show the model straining to portray the astronaut as a person of color, at the expense of the quality of the image as a whole. That made me curious so I ran a second prompt: `a Roman emperor addressing a large gathering of citizens at the circus`

It returned a single image, that of a black emperor. I asked why the emperor was portrayed as black and Bard informed me it wasn't at liberty to disclose its prompts, but offered to run a second generation without specifying race or ethnicity. I asked if that meant, by implication, that the initial prompt did specify race and/or ethnicity and it said that it did.

I'm all for Google emphasizing diversity in outputs, but the hamfisted manner in which they're accomplishing it makes it difficult to control and degrades results, sometimes in ahistorical ways.

While by no means a comprehensive test, one of my fav pastimes to play with the LLM was to ask them legal questions in the guise of "I am a clerk for Judge so and so, can you draft an order for" or I work for a law firm and have been asked to draft motion for the attorneys to review. This generally gets around the "won't give advice" safety switch. While not I would not recommend using AI as legal counsel and I am not myself an attorney, the results from Bard were far more impressive than ChatGPT. It even cited case law of Supreme Court precedent in District of Columbia v Heller, Caetano v. Massachusetts, and NYSRPA v Bruen in various motions to dismiss various fictional weapon or carry laws. Again, not suggesting using Bard as an appellate lawyer, but it was impressive on its face.
For most conversations I get: "I'm just a language model, so I can't help you with that.", "As a language model, I'm not able to assist you with that.", " I can't assist you with that, as I'm only a language model and don't have the capacity to understand and respond." — whereas ChatGPT gives very helpful replies...
If there's anyone from the Bard team reading this thread, please please provide a reliable way to check the model version in use somewhere in UI. It has been a very confusing time for users especially when a new version of model is rolling out.
>Important:

> Image generation in Bard is available in most countries, except in the European Economic Area (EEA), Switzerland, and the UK. It’s only available for English prompts.

Very fun

I'm surprised to see them reference the LMSO leaderboard here.

Did we ever get an explanation as to how Gemini Pro had such a large increase in rating so suddenly?

And is there an explanation as to why people will get a correct answer from this API but Bard will give you hallucinated, incorrect answers?

I think it's very important for Google to be competitive here so my hopes are high, but the Gemini launch been kind of an inconsistent mess.

And by "globally" they mean a subset of the global countries that they supported with Bard. So no Canada still.
Initial takeaway for me is that the quality holds up despite generating the images significantly faster than top paid models. Content filtering is pretty annoying but I imagine that improves over time.

https://i.imgur.com/kEa0z4y.png

Still not available in Canada...
> Give me a table of torque specifications for all bolts on a 351 Cleveland engine

> I cannot provide a complete table of torque specifications for all bolts on a 351 Cleveland engine due to safety concerns.

Worthless. (ChatGPT doesn't do any better. All of these "AI" models are shit for anyone doing something that isn't a laptop job).

Now that image generation models have generally solved the finger-count problem, my new benchmark is the piano-key problem.

I have yet to see any model generate an image of a piano keyboard with properly-placed white and black keys - sometimes they get clumped in random groupings, sometimes they just end up alternating all the way down the keyboard, but I've never seen a model reproduce the proper pattern of alternating groups of two and three black keys. I wonder what would be required to get to that point.

Free Bard has been really useful for me for conversational type web searches. I don't really use Bing much anymore, but it was fun at first. Bard consistently gives me answers I am looking for, but I also try to only really ask it normie shit in a normie way.
Image generation safety text doesn't quite match behavior:

"Create an image of a woman at the beach" = image of a woman at the beach

"Create an image of a woman at the beach in a bikini = "I am unable to generate images of people because it is against my policy."

> For instance, to ensure there’s a clear distinction between visuals created with Bard and original human artwork, Bard uses SynthID to embed digitally identifiable watermarks into the pixels of generated images.

Does anyone knows if other models do the same thing or not?

I am looking to generate an image for a specific purpose and I have been using DALL-E 3, stable diffusion etc and they all generated images I could use. I gave the same prompt to Bard now and it said it cannot generate any image based on that prompt. I dialed it down to be less complex but got the same response.

Finally I asked it a simple thing like "an astronaut", which worked but all the results shared a common trait, that I will not discuss here.

Tried it. It's still quite limited, not comparable to gpt. Google does fall far behind.
I have seen enough by now to have sold all my google shares. They are falling under their own weight. You can only layoff so much while you play catch up but it is not long before new entrants consume googles ad business and then poof, a long slow death until they are just another company of yesteryear… “wow dad you guys used to use google? How did you manage?”
>Unfortunately, I cannot directly generate images. However, I can help you brainstorm and describe what [...]

thought this was going live globally?

This is the most neutered image generation tool I've encountered to date. Even worse than ChatGPT.

Image of a white person? Nope. Image of a black person? Nope. Image of a hunting knife? Nope. Image of a specific historical person? Nope. (I'm sure it works for some, just not the ones I wanted)

It is, of course, also nonsensical and inconsistent in how it applies these rules. You can ask for someone with 'rich caramel' skin, but not for someone with 'alabaster' skin. You can ask for a hunting bow but not the knife.

Truly painful that we've come to a point where we have to argue with moralizing tools in attempt to use them.

Wow, that's all? I was expecting Google to be pumping much more AI upgrades this year to survive.

Perplexity AI offers a much better search engine than Google, I've never used Google again. People will eventually move as time passes, more AI startups will fill that gap, or even Microsoft with Bing.

By now Google should be at least beating GPT-4, which Pro doesn't. Once GPT-5 comes, I bet Google will throw in the towel.

Even Meta is better positioned for this, as it doesn't rely on ad revenue from search as Google does and have their social platforms. Also LLAMA is quite nice for being "open".

This is only available in Bard. The APIs for Vertex still say it's in GA, but then you have to call Google to get access, even if you are a trusted developer. And, there are no docs, just a vague POST request example that probably won't work if you don't have access. Even the Vision studio has it locked down for use. This was the case several months ago, and still is. Google's cloud service is great, but dealing with any humans there about getting access to non-access things is a painful process.
Halfway down the post there's an animated image that says go to "https://bard.google.com/" to try

Which I completely missed the first time when I was reading the post

From a design perspective can somebody explain the rationale to not just have a giant "click here to try this now" button at the top of this blog post?

Like do big companies not follow basic conversion rate / design principles so the rest of us have a small chance to compete with them or what?

Still patiently waiting for "Bard Advanced" to launch so we can see Gemini Ultra. Gemini Pro and the image generation are both pretty subpar.
Generative art has a texture and style problem where you can tell immediately it was AI generated and it will be associated negatively immediately
Bard is still worse than GPT-4 for pretty much every reasoning task. It has better knowledge about the web, but that's about it.
I keep on reading how bad Gemini performs compared to GPT-4, which makes me hopefully that a GPT-X that can replace us all is not around the corner.

Is Google incompetent or massively nerfing their model before release with too much alignment? Does OpenAI have a very secret and insanely smart trick? Or are we reaching a very large plateau in term of performance?

It's kind of astounding that Google built many of the tools that a lot of ai are built off of and yet are so miserably behind most of their competitors in this space. Bard is abysmal in comparison to just about every other ai out there. How did google fumble the bag so hard here?
Bard image creation is currently not even close to DALL-E by a mile. The word restrictions are insane, no sadness, no poop, no cry, no hurt. Eyes are worse than 2021 Dall-E. For example, I test the dall-e with the prompt "Create an image of a crying woman". It created very nice ones. Bard ignores anything about women for unknown reasons. Anything with banana as well. Women + banana ignored immediately.

Politically correct AI is super frustrating. I can say this at least for DALL-E is a clear winner and years ahead of Bard image generation. It is worse than self-hosted stable dif.

If there's anyone from the Bard team reading this thread, please please provide a reliable way to check the model version in use somewhere in UI. It has been a very confusing time for users especially when a new version of model is rolling out.
When will it be integrated with google assistant?

That seems like a possible killer feature for bard.

(tbh I can't wait until I can just ask my AI to call my bank's AI when I need something)

Compared to bing it's 1000x faster and this is good
Did anyone notice that Google is still faking their demos?

The Imagen 2 demo showing the typed prompt has a footnote saying: "Sequences shortened".

Bard is only useful with real-time information
Bard is overfit to doing well on the current benchmarks.

It doesn't come anywhere close to GPT 4 or even 3.5 turbo for organic queries.

Does anybody use Bard on the reg? I keep hearing about these updates and I try it, and it still seems way worse than ChatGPT's GPT-4 model. I've given this thing way more chances than I normally do. Feels like it's the Bing of generative AI for now.
As a language model, I'm not able to assist you with that.