Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    AI models flub these intelligence tests. Can you fare any better?

    1TB Motorola Razr Ultra (2025) drops by $700 and includes a free $100 sweetener

    Bill Gates says we’ve passed AI’s danger thresholds. Now what?

    Facebook Twitter Instagram
    • Tech
    • Gadgets
    • Spotlight
    • Gaming
    Facebook Twitter Instagram
    circuitthoughtscircuitthoughts
    Subscribe
    • Home
    • Gadgets
    • Insights
    • Apps

      Google Uses AI Searches To Detect If Someone Is In Crisis

      Gboard Magic Wand Button Will Covert Your Text To Emojis

      Android 10 & Older Devices Now Getting Automatic App Permissions Reset

      Spotify Blend Update Increases Group Sizes, Adds Celebrity Blends

      Samsung May Improve Battery Significantly With Galaxy Watch 5

    • Gear
    • Mobiles
      1. Tech
      2. Gadgets
      3. Insights
      4. View All

      AI models flub these intelligence tests. Can you fare any better?

      1TB Motorola Razr Ultra (2025) drops by $700 and includes a free $100 sweetener

      Bill Gates says we’ve passed AI’s danger thresholds. Now what?

      New Galaxy S27 Ultra renders reveal an annoying problem is about to get worse

      March Update May Have Weakened The Haptics For Pixel 6 Users

      Project 'Diamond' Is The Galaxy S23, Not A Rollable Smartphone

      The At A Glance Widget Is More Useful After March Update

      Pre-Order The OnePlus 10 Pro For Just $1 In The US

      Motorola Edge+ Review: It Checks A Lot Of Boxes

      This Smartphone Concept Design Is Different… In A Good Way

      Twitter Just Made Searching Your Direct Messages Better

      That Netflix Price Hike Is Starting To Take Place

      Latest Huawei Mobiles P50 and P50 Pro Feature Kirin Chips

      Samsung Galaxy M62 Benchmarked with Galaxy Note10’s Chipset

      9.1

      Review: T-Mobile Winning 5G Race Around the World

      8.9

      Samsung Galaxy S21 Ultra Review: the New King of Android Phones

    • Computing
    circuitthoughtscircuitthoughts
    Home»Tech»AI’s recursive self-improvement might not come so quickly after all
    Tech

    AI’s recursive self-improvement might not come so quickly after all

    adminBy No Comments3 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email

    “On the other hand, the agents were unambiguously bad at carrying out the research itself,” says Kapoor. They ran bizarre experiments (in some cases testing their hypotheses on tiny synthetic datasets), struggled to write intelligibly about their work, and made no novel contribution to their fields. “The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” he says. 

    That’s because the agents struggled to muster the creativity and judgment necessary for conducting research. They didn’t do enough to explore different ideas, and they committed to unpromising approaches too quickly. Though the agents developed novel and ambitious hypotheses resembling those that the original authors themselves started with, they rejected them on the basis of very limited data. And they couldn’t backtrack from failing approaches. They could make small pivots but could not fundamentally rethink their approach or try new ones from scratch. 

    The agents also failed to incorporate feedback from subagents or external AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They also couldn’t effectively use resources, such as tokens, compute, and time. And they couldn’t follow instructions about things like how much time to spend on different phases of the research or how long their paper could be.

    For all their failures, the agents didn’t engage in the misbehavior that researchers call “reward hacking,” hiding or misrepresenting experiments or data. Although subagents, or helper AIs that the main agent spawns to handle pieces of the work, occasionally hallucinated or misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project. 

    The reason AI models are good at research engineering but not at open-ended research may come down to how they’re trained, says Kapoor. Models get good at whatever they can be drilled on in a training regime called reinforcement learning, which is easier to apply to tasks whose success can be checked automatically. “But it’s harder to create environments to train these models when the task itself is open-ended,” he says.

    Kapoor says the team is now conducting the experiment with Mythos, Anthropic’s most advanced model, which launched in April. It was subsequently required by the Trump administration to meet various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment.

    There are some limitations to the study. It covered just two research papers, and the original authors knew the papers they were grading were generated by AI agents, which could have colored their evaluations. And the researchers had substantial discretion in designing and executing the study, meaning that their preexisting beliefs and biases could have slipped into the results. Evaluations of open-ended research trade some objectivity for a much richer test than any benchmarks can offer.

    Still, the results may temper the claims that recursive self-improvement is on the horizon. In June, Anthropic published a blog post titled “When AI Builds Itself,” charting its progress toward models that speed up their own development. In July, OpenAI advertised the fact that its new model GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work.

    #AIs #recursive #selfimprovement #quickly

    AIs quickly recursive selfimprovement
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

    Related Posts

    AI models flub these intelligence tests. Can you fare any better?

    1TB Motorola Razr Ultra (2025) drops by $700 and includes a free $100 sweetener

    Bill Gates says we’ve passed AI’s danger thresholds. Now what?

    Add A Comment

    Leave A Reply Cancel Reply

    Editors Picks
    8.5

    Apple Planning Big Mac Redesign and Half-Sized Old Mac

    Autonomous Driving Startup Attracts Chinese Investor

    Onboard Cameras Allow Disabled Quadcopters to Fly

    Top Reviews
    9.1

    Review: T-Mobile Winning 5G Race Around the World

    By
    8.9

    Xiaomi Mi 10: New Variant with Snapdragon 870 Review

    By
    8.9

    Samsung Galaxy S21 Ultra Review: the New King of Android Phones

    By
    circuitthoughts
    Facebook Twitter Instagram Pinterest Vimeo YouTube
    • Home
    • Tech
    • Gadgets
    • Mobiles
    • Our Authors
    © 2026 ThemeSphere. Designed by WPfastworld.

    Type above and press Enter to search. Press Esc to cancel.