TeleTune infers each session’s goal from usage logs and uses a text skill library to predict users’ next actions. It keeps skill edits only when they improve accuracy on held-out logs. The source says next-action accuracy on old logs tracked success in a live environment, but gives no experimental figures.
All AI updates
Oct 9
In the study, 98 patients used the research-grade diagnostic AI chatbot AMIE before their visits. Supervising doctors monitored every interaction in real time, and none had to be stopped under preset safety criteria. Clinicians said its summaries helped them prepare in 75% of cases. Larger clinical trials are still needed to assess its performance for patients.
MirroS has released AgentGarten, which connects a code-controlled physics environment to a real-time neural renderer. Agents act on visual feedback and pass written lessons to the next round. In a one-on-one hide-and-seek experiment, the hider learned to build cover by round four and the seeker to cross a wall using a ramp by round ten. The team also reported progress across four other tasks over four rounds of trials.
A post says OpenAI withdrew part of a math paper it had published two days earlier. The material does not specify which content was withdrawn.
A post describing an NVIDIA paper says VERA turns benchmark runs into more than 9,000 restartable sandboxes that score each step against real files and logs. On the reported medical research benchmark, a 9B agent scored 69.1 with both model training and skill-file edits, versus 43.3 and 56.1 with either approach alone, respectively.
VERA scores individual steps in long tasks against real files and logs, then uses those scores to choose whether to improve the model or its skill files next. On a medical research benchmark, a 9B agent scored 69.1 with both kinds of updates, compared with 56.1 for skill edits alone and 43.3 for training alone.
OpenAI released hundreds of model-generated solutions to difficult math problems this week. Researchers at the University of Cambridge and King’s College in London found at least two discrepancies between the natural-language proof and Lean code for one solution. The discrepancies do not necessarily disprove it, but the researchers say such proofs still need human scrutiny.
Oct 8
PaperBenchX turns 93 papers across 12 research areas into reproduction tasks, requiring AI agents to submit workflows that can be rerun in an isolated environment. In this benchmark, the best-performing GPT-6 Astra fully reproduced 13.98% of tasks; the team has opened 12 tasks and kept 81 closed.
PivotOPD identifies task disrupting actions during training and teaches agents to recover after making them. In replays of 72 pivotal mistakes, it recovered 72.7% of the time, versus 20.3% for standard on-policy distillation. The method depends on replayable environments, and its code has not yet been released.
A new paper describes a chatbot silently correcting an error while translating an incorrect math proof into Lean, producing a valid Lean proof that does not establish the correctness of the original natural language proof. The accompanying account also says deciding whether a statement can be translated faithfully is harder than the Halting problem, so no AI translator can always do it.
NAT-ARC pretrains a visual encoder on ImageNet images, trains it on ARC grid puzzles, and fine-tunes it separately for each puzzle at test time. The reported ARC-1 pass@2 scores are 63.4±0.7% for the best single model and 70.2±0.6% for an ensemble that combines models trained with different pretraining strategies.
Researchers introduce Normalizing Trajectory Models (NTM), which models each image-generation step as a conditional normalizing flow and uses exact likelihood training. The paper reports that NTM matches or outperforms strong baselines on text-to-image benchmarks in four sampling steps while retaining exact likelihood over the generative trajectory; the supplied material gives no scores.
Roughly a third of the sky has never been observed in ultraviolet light. The researcher combined existing observations and used Claude Science to estimate the gaps from visible, infrared and radio data; the map distinguishes measured from predicted areas and provides uncertainty estimates.
A team led by Jiang Zhongyi and He Guangwei at Tianjin University embedded single-crystal COF in a polymer membrane, reaching a 75.5% filler volume fraction. In simulated flue gas mixture tests, the membrane achieved CO₂ permeability of 74,800 Barrer and CO₂/N₂ selectivity of about 20; a preliminary techno-economic analysis estimated capture costs at about $38 per tonne of CO₂ under specified process and cost assumptions.
Oct 7
QbitAI reports that the Royal Swedish Academy of Sciences awarded the 2026 Nobel Prize in Chemistry to Henri B. Kagan and Kenso Soai for their discoveries of nonlinear effects and autocatalysis in asymmetric organic synthesis. It says Soai's 1995 experiment amplified one mirror-image molecule's initial 2% lead to 87%.
The paper proposes JAZ, which lets an agent access its prompt and full history as variables and pass them to subagents. On StuLife tasks requiring facts from more than 50 tasks earlier, the source reports a 69.9% pass rate for JAZ versus 61.8% for Letta, at less than half the cost. These are reported results.
Looped language models pass each token through the same layers multiple times, increasing cache use. The paper proposes keeping the latest 128 tokens exact and compressing older tokens into small vectors that each loop can read directly. On Ouro models at 16K tokens, throughput rose by up to 7.4 times, while accuracy on math, knowledge, and reasoning stayed above 97% of the original.
OpenAI released 722 mathematics manuscripts from an unreleased internal model, organized into 372 groups, along with papers, source code and some Lean proofs. They address questions including a zero-free region for the Riemann zeta function and specific cases of the Hodge conjecture. The repository says verification varies across results and some lack Lean formal proofs; an independent advisory group says offering advice does not amount to endorsing the results.
展开进展与来源
A post says OpenAI released 722 AI-generated math proofs, with related papers covering 372 mathematical results. One result lowers the proven bound for multiplying two n×n matrices from about n^2.371177 to about n^2.25 and includes a Lean formalization. It remains a theoretical bound and does not yet provide a faster routine for real GPU workloads.
同一事件,精选展示《OpenAI公布未发布模型生成的数学研究成果》展开进展与来源
The release contains 722 manuscripts covering 372 families of results. Advisory group AGMAI says it includes answers to “hundreds” of open problems. The materials also include some reasoning summaries and compute estimates, but mathematicians still need to assess the work, and its full impact remains unclear.