artificial
artificial

I ran GPT-6 Astra against 7 real signup CAPTCHAs

OpenAI put GPT-6 Astra into the API today, so now anyone can actually call it instead of waiting for the special-access rollout. Good timing too, since everyone's still reposting Sharif Shameem's video from a couple days ago, where GPT-6 …

¿Por qué medimos el progreso de la IA por lo que puede reemplazar, en lugar de por lo que permite hacer a los humanos?

¿Por qué planteamos cada nueva capacidad de IA en términos de lo que puede reemplazar? Cada vez que la IA adquiere una nueva capacidad, surge casi de inmediato una pregunta familiar: ¿Qué reemplazará? Si un modelo mejora en programación, nos preguntamo…

We need free market and foreign AI models to keep companies competitive

There’s a lot of talk from different AI companies that are pushing the narrative that AI is simply too dangerous to be distributed without severe regulations and reviews. While I agree with this, I have to point out that the solution isn’t shutting dow…

VSArena v0.6.0 — a new Studio for running and inspecting embodied AI policies in the browser

I just released VSArena v0.6.0, a major update to the browser-based Studio for VSArena. VSArena is an open evaluation arena for Vision-Language-Action (VLA) and embodied AI policies, built around browser-native 3D physics. The goal is simple: mak…

Notion ai finally clicked for me once i stopped using it as a writer

for months i thought it was overpriced because i kept asking it to write stuff and the writing was kinda mid. turns out that's not the point. the point is asking it questions about your own pages. "what did i say i'd do about X last week&q…

Anthropic Researcher Abruptly Resigns Before Warning That AI ‘Could Kill Us All By The End Of The Decade’ In Alarming Rant

submitted by /u/ComicSandsNews [link] [comments]

The ‘AI Will Kill Us’ Claim is Just a Wienie

On this week's episode of Lanterns (it's a great show about Green Lanterns), John Stewart mentions wienies. It's not what you think. Walt Disney came up with the idea. It's an architectural feature that can draw your attention aw…

i built a benchmark to test whether LLMs can understand and create jokes

lolbench – LLMs take three tests: – explain why jokes work (or don't) – write jokes under shared premises – predict which jokes humans prefer The finding so far that surprised me: every model aces explaining real jokes (95%+) but drops hard on ex…

Is there a way to get AI to do a group chat thing? Like you bouncing ideas off of them, but instead of 1 LLM it is multiple that can refine that idea

So sometimes when I try to understand something or think about something. I might use a LLM like Gemini or Grok. Like for example, I'm thinking about how water down the cyber security should be on my main email account. Like if I should just do a s…

Ken Cox never fired anyone. His headcount fell from 175 to 3 anyway — here’s the rule that did it.

TL;DR: Ken Cox never fired anyone. His headcount just kept dropping anyway. 175 down to 3 is the actual number — and the reason is one operating rule he'd been running for years before AI made it urgent: if it can be automated, it will be a…