Toggle light / dark theme

Get the latest international news and world events from around the world.

Log in for authorized contributors

Tail-Likelihood Reinforcement Learning: Teaching AI to Keep Its Best Possibilities Alive

Most reinforcement learning methods train an AI to maximize its average reward. That sounds sensible—but averages can hide something important. https://arxiv.org/abs/2609.

Think of two AI systems that both score an average of 70. One almost always produces results around 70. The other usually scores 60, but occasionally discovers an exceptional solution scoring 100. If we are allowed to sample the AI many times and choose the best result, the second system could be far more valuable.

This is the problem addressed by Tail-Likelihood Reinforcement Learning (TailRL).

Instead of asking only, “What reward does the AI achieve on average?”, TailRL asks a different question across the entire range of outcomes:

“How likely is the AI to produce something better than a particular reward threshold?”

The method effectively turns continuous rewards into a collection of success-or-failure questions. Importantly, it gives greater learning weight to rare, high-reward outcomes rather than allowing them to disappear into the average.

There is an interesting connection to Best-of-k sampling: when we generate multiple attempts and keep the best one, performance depends on whether the model has retained enough probability mass around those unusually good solutions. TailRL explicitly trains the model to preserve that potential.

The practical advantage is significant. The researchers report that TailRL helps models avoid getting trapped in mediocre solutions across tasks including object localization, maze navigation, GUI grounding, and code optimization. More importantly, the resulting models become better at exploiting additional sampling at inference time—the more attempts you give them, the more useful those attempts become.

Coolest lava world yet with signs of an atmosphere offers clues to early Earth

In the search for extraterrestrial life, it makes sense to first look for rocky planets with an atmosphere, like Earth. Without an atmosphere, a planet can’t have surface water. But of the more than 6,300 exoplanets cataloged thus far, the vast majority are not rocky, and only a handful of the rocky worlds appear to have an atmosphere.

In a study published in The Astrophysical Journal Letters, a group led by University of Chicago scientist Brandon Park Coy reports another: a rocky super-Earth 154 light-years away in the constellation Pisces. Named HD 3,167 b, this very hot “lava world” zips around its host star in just one Earth day.

“What’s so surprising is that the closer a rocky planet orbits its star, the harder it should be to have an atmosphere, because it’s bombarded by stellar wind and gets more high-energy photons from the star. But it seems that many of these lava worlds do,” explained Edwin Kite, UChicago associate professor of geophysical sciences and co-author of the study. “These planets are too hot for life, but by studying them, we can say something about the processes that matter for other rocky worlds.”

10XMe · AI that actually works for you

Ever wondered what actually makes an AI “agent” different from a basic chatbot? 🤖💡

While standard chatbots just generate text, a true AI Agent can think, plan, use tools, and execute complex workflows autonomously (or alongside humans).

Here is a simple breakdown of how an AI Agent works under the hood:

🧠 The Brain (Reasoning & Memory): It uses LLMs to process information, leveraging short-term context and long-term vector memory to recall instructions and past interactions.

🎯 Planning & Decisions: Before acting, it breaks big goals down into smaller, actionable steps. If it can answer directly, it does; if not, it triggers the action layer!

🛠️ Tools & Execution: It interacts with real-world applications—running code, querying databases, triggering API calls, and reading files to get actual work done.

🛡️️ Safety & Monitoring: Surrounding the whole process are strict guardrails (human-in-the-loop checkpoints, rate limits, content filtering) and observability tools to track accuracy, costs, and performance.

A Case of Coffee Enema Induced Rectal Burn, Proctocolitis, Sub-Acute Intestinal Obstruction, and Electrolyte Imbalance A Case Report

Abstract

Coffee enemas are a form of complementary and alternative medicine (CAM) promoted by some practitioners and celebrities without credible scientific evidence. Their use is associated with several adverse effects, some of which can be life-threatening, including rectal burns, proctocolitis, intestinal obstruction, and electrolyte imbalances, which are well-documented but under reported. Only a handful of comparable cases exist in the literature, highlighting the need for greater clinician awareness and patient education. We present the case of a 30-year-old obese male with a recent diagnosis of T2 diabetes mellitus who developed severe proctocolitis, a fibrotic stricture leading to subacute intestinal obstruction, and a significant electrolyte imbalance following a self-administered hot coffee enema. The patient required intensive care management and partial colectomy was planned. This case report shows that coffee enema sometimes can cause serious complications. Healthcare providers and patients should be aware of the risk and should use coffee enema with caution.

Coffee Enema, Rectal Burn, Proctocolitis, Fibrotic Stricture, Subacute Intestinal Obstruction, Electrolyte Imbalance, Diabetes Mellitus, Alternative Medicine, Clinician Awareness, Patient Education

Brain circuit remembers stresses to help shape reactions to them

Researchers at the Icahn School of Medicine at Mount Sinai have identified a previously overlooked brain circuit that helps explain how prior adversity can make the brain more reactive to future stress.

The study, published Sept. 30 in Nature and conducted in mice, identifies a small, deep brain structure called the anterior hypothalamic nucleus (AHN) as a critical hub that scales the brain’s response to threatening events. The research team found that dialing its activity up or down directly altered how strongly the animals responded to stress.

“Why do some people develop debilitating mental health conditions in response to stress while others do not? One known risk factor for heightened stress sensitivity is a history of prior stress, such as early childhood adversity or adult traumatic stress. However, at a biological level, we still do not fully understand why this is the case,” says Zachary Pennington, Ph.D., lead author of the paper, who conducted the research as a postdoctoral fellow in the Cai Lab at Mount Sinai and is now an assistant professor of psychology and a member of the Djavad Mowafaghian Centre for Brain Health at the University of British Columbia.

U.S. to lend $4.2 billion to Vistra to boost nuclear power output, source says

U.S. Energy Secretary Chris Wright will announce the loan on Monday at Vistra’s nuclear plant along Lake Erie in Ohio, the person said, speaking on condition of anonymity.

The loan would be used to boost ⁠power ‌output, or “uprate” at least three of Vistra’s ⁠four nuclear stations, which can be done without requiring new licenses from the Nuclear Regulatory Commission.

The Department of Energy and Vistra did not immediately respond to ‌requests for comment.

Google Suncatcher’s Next Act: 2 Satellites by 2027

Planet built its business on Earth-imaging satellites, not AI accelerators, which makes the pairing worth a second look. The company’s own statement on the deal frames it less as a subcontracting job and more as co-ownership of the research question. Planet said it was pleased to announce its participation in Google’s Project Suncatcher, calling it a bold research initiative to explore building scalable machine learning compute systems in space, according to Planet’s own announcement.

Planet brings something Google doesn’t have in-house at scale: a working satellite bus manufacturing line and an existing relationship with launch providers. Google supplies the chips and the machine learning workload. Planet supplies the spacecraft, the power system, and the operational experience of actually running satellites day to day. Satellite imaging companies like Planet already solved a version of this problem, keeping sensors cool, powered, and pointed correctly in orbit, which is a large part of what an AI-compute satellite also needs. That division of labor is probably why the partnership extends into the 2027 mission rather than ending after MVP.

The hardest unsolved problem in orbital computing isn’t the chips, it’s the wiring. Training large models on Earth depends on extremely fast links between GPUs and TPUs sitting centimeters apart in the same rack. Spread that same cluster across satellites orbiting hundreds of kilometers apart, and the interconnect problem changes entirely. Google’s answer is optical: the satellites will communicate via lasers, according to Google’s fact sheet, rather than conventional radio-frequency links.

/* */