• About
  • Advertise
  • Privacy Policy
  • Contact
Over View - Your Daily News Source
  • Home
  • News
    • Business
    • Politics
    • Science
  • Lifestyle
    • Food
    • Travel
    • Health
    • Fashion
  • Entertainment
    • Entertainment
    • Sports
  • Tech
No Result
View All Result
  • Home
  • News
    • Business
    • Politics
    • Science
  • Lifestyle
    • Food
    • Travel
    • Health
    • Fashion
  • Entertainment
    • Entertainment
    • Sports
  • Tech
No Result
View All Result
Over View - Your Daily News Source
No Result
View All Result
Home News Science

Google Gemini 3.8 Flash Has a Mix of Improved Benchmarks

admin by admin
September 2, 2026
in Science
0
Google Gemini 3.8 Flash Has a Mix of Improved Benchmarks
0
SHARES
8
VIEWS

Gemini 3.8 Flash is Google’s newest Flash-tier model released September 2, 2026. It is the most capable Flash model yet for mid-difficulty long-horizon software engineering, autonomous agents, and enterprise workflows, while keeping Flash-level speed and the same introductory pricing as 3.7 Flash. It is a value model.

Pricing is the same as 3.7 Flash

Introductory (through Dec 31, 2026): $0.75 / 1M input, $3.75 / 1M output
Regular (from Jan 1, 2027): $1.50 / $7.50

Where it is strong
Vals Finance Agent v2 — 61.4% (best)
Harvey’s Legal Agent Benchmark — 10.0% (best. all models score poorly here)
Terminal-bench 2.1 — 89.4% (narrowly best)
CharXiv Reasoning (charts, no tools) — 86.2% (best)
LVBench long video — 87.8% agentic / 87.1% static (best)
HLE-Verified — 54.9% (best)
BioMysteryBench (Human Difficult) — 56.5% (best)
LABBench2 — 86.2% (best)

These wins cluster in multimodal understanding (video + charts), domain-specific professional agents (finance, legal, biology), and mid-difficulty terminal/agentic coding.

Where it is behind

DeepSWE v1 (hard long-horizon SWE) — 71.0% vs Opus 5’s 74.0% and GPT-5.6 Sol’s 72.7%
GDPVal-AA v2 knowledge-work Elo — 1545 vs Opus 5’s 1824
Terminal-bench 4.0 (harder general agents) — 19.1% vs Opus 5’s 51.8%
OSWorld-2.0 computer use — 59.0% vs Opus 5’s 75.4%
GDP.PDF — 35.0% vs GPT-5.6 Sol’s 40.0%

Brian Wang is a Futurist Thought Leader and a popular Science blogger with 1 million readers per month. His blog Nextbigfuture.com is ranked #1 Science News Blog. It covers many disruptive technology and trends including Space, Robotics, Artificial Intelligence, Medicine, Anti-aging Biotechnology, and Nanotechnology.

Known for identifying cutting edge technologies, he is currently a Co-Founder of a startup and fundraiser for high potential early-stage companies. He is the Head of Research for Allocations for deep technology investments and an Angel Investor at Space Angels.

A frequent speaker at corporations, he has been a TEDx speaker, a Singularity University speaker and guest at numerous interviews for radio and podcasts.  He is open to public speaking and advising engagements.

Read More

Previous Post

“Broken heart syndrome” can look just like a heart attack. This test can tell them apart

Next Post

Dana White suggests ‘common sense’ rule change after Paddy Pimblett’s UFC 329 fine causes uproar

Next Post
Dana White suggests ‘common sense’ rule change after Paddy Pimblett’s UFC 329 fine causes uproar

Dana White suggests ‘common sense’ rule change after Paddy Pimblett’s UFC 329 fine causes uproar

  • About
  • Advertise
  • Privacy Policy
  • Contact

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.

No Result
View All Result
  • Entertainment
    • Entertainment
    • Sports
  • Lifestyle
    • Fashion
    • Health
    • Travel
    • Food
  • News
    • Business
    • Politics
    • Science
  • Tech

© 2026 JNews - Premium WordPress news & magazine theme by Jegtheme.