🔗 원문 전체 보기 → Venturesquare.net

요약
AI의 역할이 과학 논문 검색·요약에서 가설 생성과 추론, 실험 설계·검증까지 확대되고 있다. 국내 아스테로모프는 범용 AI에 과학 데이터와 전문 도구, 다중 에이전트 검증을 결합하는 추론 레이어를 개발하고 있다. 라일라 사이언스와 피리오딕 랩스는 AI와 실제 실험 환경을 연결해 실험 결과가 다시 모델과 다음 연구에 반영되는 구조를 구축하고 있다. 과학 AI 경쟁의 무게중심도 모델 성능에서 추론과 실제 실험을 연결하는 연구 시스템으로 이동하고 있다.
본문
- 아스테로모프·라일라 사이언스·피리오딕 랩스 등 AI 추론과 실험 연결…과학 연구 자동화 경쟁
- 범용 AI에 전문 데이터·연구도구·자율실험실 결합…검색·요약 넘어 가설 생성과 검증으로 확장
과학 연구에서 AI의 역할이 논문을 찾고 요약하는 보조도구에서 연구 과정에 직접 참여하는 단계로 이동하고 있다. 복잡한 문제를 여러 경로로 추론하고 가설을 세운 뒤 실험 결과를 다시 다음 추론에 반영하는 방식이다. 특히 생물학과 신소재처럼 실제 실험을 거쳐야 가설을 검증할 수 있는 분야에서는 AI 모델 자체의 성능뿐 아니라 전문 데이터와 연구 도구, 실험 환경을 어떻게 연결하느냐가 중요해지고 있다. 범용 AI와 실제 연구 사이를 연결하는 ‘Scientific AI(과학 AI)’가 새로운 기술 경쟁 영역으로 떠오르는 배경이다.
최근 수학 분야에서는 AI가 대규모 추론과 검증 과정에 투입되는 사례가 잇따랐다. 오픈AI는 이달 내부 연구 모델을 사용하는 약 1만개의 동시 실행 에이전트를 투입해 나비에-스토크스 밀레니엄 문제에 대한 해법을 제안했다. 에이전트가 해법에 도달하기까지 약 88시간이 걸렸고 이후 GPT-6 Astra를 이용한 Lean 형식화와 검증에 17시간을 추가로 사용했다.

앤트로픽은 클로드를 이용해 페르마의 마지막 정리의 기존 증명을 11일 동안 Lean으로 형식화했다. 이는 새로운 증명을 찾아낸 사례라기보다 기존 수학적 증명을 컴퓨터가 처음부터 끝까지 검증할 수 있는 형태로 전환한 것으로, 최종 결과물에는 약 1300만줄의 Lean 코드가 사용됐다.
수학에서는 형식 검증을 통해 결과를 확인할 수 있지만 생물학이나 소재 연구는 상황이 다르다. AI가 가능성이 높은 가설을 제시하더라도 실제 실험을 통해 결과를 확인하고 실패한 가설을 수정하면서 다음 연구 방향을 정해야 한다. 이에 따라 AI의 추론 능력을 전문 지식과 데이터, 연구 도구, 실험·검증 과정까지 연결하는 시스템 개발이 이어지고 있다.
국내에서는 아스테로모프가 범용 AI 모델과 실제 과학 연구 사이를 연결하는 ‘추론 레이어(reasoning layer)’를 개발하고 있다. 대규모언어모델(LLM)에 분야별 과학 데이터와 전문 도구, 여러 AI 에이전트의 협업·검증 과정을 결합해 반복적인 추론을 거치면서 가설을 발전시키는 구조다.
개념검증(PoC) 시스템 ‘틸라(Thyla) 1.0’은 특정 기반 모델에 종속되지 않는 모델 애그노스틱 구조를 적용했다. 문제에 따라 여러 AI 에이전트와 과학 데이터, 전문 도구를 조합한다.
아스테로모프 자체 평가에서 기반 모델로 사용한 GLM-5.1은 ‘Frontier Science-Research’에서 10.0%를 기록했고 추론 레이어를 적용한 틸라 1.0은 33.9%를 기록했다. 회사는 이를 기반 모델의 성능뿐 아니라 모델 위에서 추론과 검증을 어떻게 구성하는지가 과학 문제 해결 성능에 영향을 줄 수 있다는 결과로 해석했다.
지난 7월 열린 제56회 국제물리올림피아드(IPhO 2026) 이론시험을 이용한 평가도 진행했다. 회사에 따르면 자체 AI 시스템은 실제 참가자와 같은 5시간 동안 외부 AI 모델이나 인터넷에 접속하지 않고 추론과 검증, 답안 작성을 수행해 30점 만점에 28.6점을 받았다.
가설→실험→데이터→다음 가설…AI 연구가 ‘폐쇄루프’로
미국 라일라 사이언스(Lila Sciences)는 AI 모델과 로봇 기반 자율실험실을 결합한 ‘AI Science Factory’를 구축하고 있다. AI가 가설을 만들고 실험을 설계하면 자동화된 실험 환경에서 이를 수행하고 결과 데이터를 다시 모델에 넣어 다음 실험을 설계하는 폐쇄루프 방식이다.
라일라는 이 시스템을 이용해 95만개 이상의 mRNA 서열을 실제 합성하고 살아있는 세포에서 측정한 데이터를 AI 모델 학습에 활용했다. 이후 해당 플랫폼을 체내 CAR-T 연구에 적용했으며 회사가 공개한 비인간 영장류 실험에서는 비교 대상으로 설정한 조성보다 더 깊고 지속적인 B세포 제거 결과를 확인했다고 설명했다.
피리오딕 랩스(Periodic Labs)는 AI와 자율실험실을 결합해 초전도체와 자성 소재 등 새로운 물질을 탐색한다. 실험 결과를 AI가 분석해 다음 실험에 반영하고 새로 만들어진 데이터를 다시 모델 학습에 사용하는 구조다.
회사는 지난 15일 자체 실험실 데이터를 학습에 활용한 1조 파라미터 규모의 ‘Periodic Neon’을 공개했다. 이 모델은 X선 회절(XRD) 데이터를 분석해 실험에서 생성된 물질을 판별하는 데 초점을 맞추고 있으며 실제 실험실에 배치됐다.
디스커버리 루프(Discovery Loop)는 특정 과학 분야보다 연구와 엔지니어링 과정 자체를 AI로 자동화하는 데 초점을 맞춘다. 연구자가 반복해 수행하는 실험 제안과 실행, 결과 분석, 후속 실험 설계를 AI가 연속적으로 처리하고 다수의 실험을 병렬로 수행해 다음 연구 방향을 찾는 구조다. 머신러닝 연구와 엔지니어링을 시작으로 의약품 개발과 의료정보학, 태양에너지, 청정수 등으로 적용 분야를 확대한다는 구상이다.
이들 기업의 접근법은 서로 다르지만 AI를 연구자의 질문에 답하는 도구에서 ‘가설-실험-검증-수정’의 반복 과정에 참여하는 시스템으로 확장한다는 점에서는 맞닿아 있다. AI 모델의 추론 성능 경쟁이 실제 실험 데이터와 연구 장비, 전문 도구를 얼마나 효과적으로 연결할 수 있는지를 겨루는 단계로 이동하는 셈이다.
아스테로모프 관계자는 “향후 과학 AI의 경쟁력은 단순히 벤치마크에서 높은 점수를 내는 것을 넘어 실제 연구 현장에서 의미 있는 가설을 제시하고 이를 실험을 통해 검증하는 과정까지 얼마나 효과적으로 이어갈 수 있느냐에 달려 있다”며 “AI의 추론과 실제 실험을 긴밀하게 연결해 복잡한 과학 연구를 반복 가능한 체계로 만드는 데 집중하겠다”고 말했다.
Like this:
Like Loading...Related
Asteromorph, AI that used to search for papers even formulates hypotheses and conducts experiments… The 'Scientific AI' competition
The role of AI in scientific research is shifting from an auxiliary tool for finding and summarizing papers to a stage where it directly participates in the research process. This approach involves reasoning through multiple paths to solve complex problems, formulating hypotheses, and then incorporating experimental results into subsequent inferences. Particularly in fields like biology and new materials science, where hypothesis verification requires actual experiments, the ability to connect specialized data, research tools, and experimental environments is becoming increasingly important, in addition to the performance of the AI model itself. This is the backdrop against which 'Scientific AI,' which bridges general-purpose AI with actual research, is emerging as a new area of technological competition.
Recently, there have been a series of instances in the field of mathematics where AI has been deployed for large-scale inference and verification processes. This month, OpenAI proposed a solution to the Navier-Stokes Millennium Problem by deploying approximately 10,000 concurrent agents using an internal research model. It took the agents about 88 hours to reach the solution, followed by an additional 17 hours for Lean formalization and verification using GPT-6 Astra.

Antropic used Claude to format the existing proof of Fermat's Last Theorem into Lean over 11 days. This was not so much a case of finding a new proof, but rather converting an existing mathematical proof into a form that a computer could verify from start to finish, and the final result used about 13 million lines of Lean code.
In mathematics, results can be verified through formal testing, but the situation is different in biology or materials research. Even if AI proposes a highly probable hypothesis, the direction of future research must be determined by verifying results through actual experiments and modifying failed hypotheses. Accordingly, the development of systems that connect AI's reasoning capabilities with specialized knowledge, data, research tools, and experimental and verification processes is ongoing.
In Korea, Asteromorph is developing a 'reasoning layer' that bridges general-purpose AI models with actual scientific research. It is a structure that develops hypotheses through iterative reasoning by combining Large-Scale Language Models (LLMs) with domain-specific scientific data, specialized tools, and the collaboration and verification processes of multiple AI agents.
The proof-of-concept (PoC) system 'Thyla 1.0' applies a model-agnostic structure that is not dependent on a specific base model. It combines various AI agents, scientific data, and expert tools depending on the problem.
In Asteromorph's internal evaluation, GLM-5.1, used as the base model, scored 10.0% on 'Frontier Science-Research,' while Thila 1.0, with an inference layer applied, scored 33.9%. The company interpreted this result as indicating that not only the performance of the base model but also how inference and validation are structured on top of the model can affect scientific problem-solving performance.
An evaluation was also conducted using the theory exam from the 56th International Physics Olympiad (IPhO 2026), held last July. According to the company, its proprietary AI system performed reasoning, verification, and answer generation for five hours—the same duration as actual participants—without accessing external AI models or the internet, and received a score of 28.6 out of 30.
Hypothesis → Experiment → Data → Next Hypothesis… AI Research Becomes a 'Closed Loop'
Lila Sciences, a U.S. company, is building an 'AI Science Factory' that combines AI models with robot-based autonomous laboratories. It operates on a closed-loop system where AI generates hypotheses and designs experiments, executes them in an automated environment, and feeds the resulting data back into the model to design the next experiment.
Layla used this system to synthesize over 950,000 mRNA sequences and utilized data measured in living cells to train an AI model. Subsequently, the platform was applied to in vivo CAR-T research, and the company explained that in non-human primate experiments disclosed by the company, results showed deeper and more sustained B-cell removal compared to the composition set as a comparison.
Periodic Labs combines AI with autonomous laboratories to explore new materials, such as superconductors and magnetic materials. The structure involves AI analyzing experimental results to inform subsequent experiments, and using the newly generated data to train models.
On the 15th, the company unveiled 'Periodic Neon,' a model with 1 trillion parameters trained on its own laboratory data. This model focuses on identifying substances generated in experiments by analyzing X-ray diffraction (XRD) data and has been deployed in actual laboratories.
The Discovery Loop focuses on automating the research and engineering processes themselves using AI, rather than on specific scientific fields. It is a structure in which AI continuously processes the repetitive tasks of researchers—such as proposing and executing experiments, analyzing results, and designing subsequent experiments—and performs multiple experiments in parallel to identify the next direction of research. The plan is to expand application fields, starting with machine learning research and engineering, to areas such as drug development, medical informatics, solar energy, and clean water.
Although these companies differ in their approaches, they are aligned in that they expand AI from a tool that answers researchers' questions to a system that participates in the iterative process of 'hypothesis-experiment-verification.' In essence, the competition for AI model inference performance is shifting to a stage where it competes on how effectively actual experimental data, research equipment, and specialized tools can be connected.
An official from Asteromorph stated, “The future competitiveness of scientific AI depends not merely on achieving high scores on benchmarks, but on how effectively it can extend the process from proposing meaningful hypotheses in actual research settings to verifying them through experiments,” adding, “We will focus on closely linking AI reasoning with actual experiments to transform complex scientific research into a repeatable system.”
アステロモフ、論文探していたAIが仮説を立てて実験まで… 「Scientific AI」競争
科学研究におけるAIの役割が論文を探し、要約する補助ツールから研究過程に直接参加する段階に移行している。複雑な問題を複数の経路で推論し、仮説を立てた後、実験結果を再び次の推論に反映する方式だ。特に生物学や新素材のように実際の実験を経て仮説を検証できる分野では、AIモデル自体の性能だけでなく、専門データと研究ツール、実験環境をどのように連結するかが重要になっている。汎用AIと実際の研究を結ぶ「Scientific AI(科学AI)」が新しい技術競争領域として浮上する背景だ。
最近、数学分野ではAIが大規模な推論と検証過程に投入される事例が相次いだ。 OpenAIは今月、内部研究モデルを使用する約1万個の同時実行エージェントを投入し、ナビエ・ストークス・ミレニアム問題に対する解決策を提案した。エージェントが解決に到達するまでに約88時間かかり、その後GPT-6 Astraを利用したLean形式化と検証に17時間をさらに使用した。

アントロピックはクロードを利用してペルマの最後の整理の既存の証明を11日間Leanに形式化した。これは新しい証明を見つけた事例よりも既存の数学的証明をコンピュータが最初から最後まで検証できる形に切り替えたもので、最終結果物には約1300万ジュールのLeanコードが使われた。
数学では型式検証で結果を確認することができますが、生物学や素材研究は状況が異なります。 AIが可能性の高い仮説を提示しても、実際の実験を通じて結果を確認し、失敗した仮説を修正しながら、次の研究方向を定めるべきである。これにより、AIの推論能力を専門知識とデータ、研究ツール、実験・検証過程まで連結するシステム開発が続いている。
国内ではアステロモルフが汎用AIモデルと実際の科学研究との間を結ぶ「推論レイヤー(reasoning layer)」を開発している。大規模言語モデル(LLM)に分野別科学データと専門ツール、複数のAIエージェントの協業・検証過程を組み合わせて繰り返し推論を経て仮説を発展させる仕組みだ。
概念検証(PoC)システム「Thyla 1.0」は、特定のベースモデルに依存しないモデルアグノスティック構造を適用した。問題に応じて、複数のAIエージェントと科学データ、専門ツールを組み合わせます。
アステロモルフの自己評価でベースモデルとして使用したGLM-5.1は「Frontier Science-Research」で10.0%を記録し、推論層を適用したティラ1.0は33.9%を記録した。同社は、これを基盤とするモデルの性能だけでなく、モデル上で推論と検証をどのように構成するかが科学問題解決の性能に影響を与える可能性があるという結果と解釈した。
去る7月に開かれた第56回国際物理オリンピアード(IPhO 2026)理論試験を利用した評価も行った。同社によると、独自のAIシステムは、実際の参加者と同じ5時間、外部AIモデルやインターネットに接続せずに推論と検証、答案作成を行い、30点満点で28.6点を受けた。
仮説→実験→データ→次の仮説… AI研究が「閉ループ」として
米国ライラサイエンス(Lila Sciences)は、AIモデルとロボットベースの自律実験室を組み合わせた「AI Science Factory」を構築している。 AIが仮説を作成して実験を設計すると、自動化された実験環境でこれを行い、結果データを再モデルに入れて次の実験を設計する閉ループ方式だ。
ライラはこのシステムを用いて95万個以上のmRNA配列を実際に合成し、生きている細胞で測定したデータをAIモデル学習に活用した。その後、当該プラットフォームを体内CAR-T研究に適用し、会社が公開した非ヒト霊長類実験では、比較対象に設定した組成よりも深く持続的なB細胞除去の結果を確認したと説明した。
ピリオディックラプス(Periodic Labs)はAIと自律実験室を組み合わせて超伝導体や磁性材料などの新しい物質を探索する。実験結果をAIが分析し、次の実験に反映し、新しく作成されたデータを再びモデル学習に使用する仕組みだ。
同社は15日、独自の実験室データを学習に活用した1兆パラメータ規模の「Periodic Neon」を公開した。このモデルは、X線回折(XRD)データを分析して実験で生成された物質を判別することに焦点を当てており、実際の実験室に配置された。
ディスカバリーループは、特定の科学分野よりも研究とエンジニアリングプロセス自体をAIで自動化することに焦点を当てています。研究者が繰り返し行う実験提案と実行、結果分析、後続の実験設計をAIが連続的に処理し、多数の実験を並列に行い、次の研究方向を探す仕組みだ。機械学習研究とエンジニアリングをはじめ、医薬品開発と医療情報学、太陽エネルギー、清浄水などで適用分野を拡大するという構想だ。
これらの企業のアプローチは互いに異なっているが、AIを研究者の質問に答えるツールで「仮説-実験-検証-修正]の繰り返し過程に参加するシステムに拡張するという点では接している。 AIモデルの推論性能競争が実際の実験データと研究機器、専門ツールをどれだけ効果的に接続できるかを競う段階に移行するわけだ。
アステロモフ関係者は「今後の科学AIの競争力は、単にベンチマークで高いスコアを出すことを超えて、実際の研究現場で意味のある仮説を提示し、これを実験を通じて検証する過程まで、どれだけ効果的につながることができるかにかかっている」とし「AIの推論と実際の実験を緊密に結びつけ、複雑な科学研究を繰り返す」
Asteromorph,一款曾经用于搜索论文甚至能够提出假设并进行实验的人工智能……“科学人工智能”竞赛
人工智能在科学研究中的角色正从查找和总结论文的辅助工具,转变为直接参与研究过程。这种方法涉及通过多路径推理来解决复杂问题,提出假设,并将实验结果融入后续推论。尤其是在生物学和新材料科学等领域,假设验证需要实际实验,因此,除了人工智能模型本身的性能之外,连接专业数据、研究工具和实验环境的能力也变得日益重要。正是在这样的背景下,将通用人工智能与实际研究相结合的“科学人工智能”正在崛起,成为新的技术竞争领域。
近期,人工智能在数学领域被应用于大规模推理和验证的案例层出不穷。本月,OpenAI 利用其内部研究模型,部署了约 10,000 个并发智能体,提出了解决纳维-斯托克斯千年难题的方案。这些智能体耗时约 88 小时得出解决方案,随后又花费 17 小时使用 GPT-6 Astra 进行精益形式化和验证。

Antropic 使用 Claude 在 11 天内将费马大定理的现有证明格式化为 Lean 代码。这与其说是寻找新的证明,不如说是将现有的数学证明转换为计算机可以从头到尾验证的形式,最终结果使用了大约 1300 万行 Lean 代码。
在数学领域,结果可以通过正式测试进行验证,但在生物学或材料研究中情况则有所不同。即使人工智能提出了一个极有可能成立的假设,未来的研究方向也必须通过实际实验验证结果并修正失败的假设来确定。因此,将人工智能的推理能力与专业知识、数据、研究工具以及实验和验证流程相结合的系统开发工作仍在进行中。
在韩国,Asteromorph公司正在开发一种“推理层”,旨在将通用人工智能模型与实际科学研究连接起来。该架构通过迭代推理构建假设,其方法是将大规模语言模型(LLM)与特定领域的科学数据、专用工具以及多个人工智能代理的协作和验证过程相结合。
概念验证(PoC)系统“Thyla 1.0”采用了一种与模型无关的结构,不依赖于特定的基础模型。它根据问题的不同,结合了各种人工智能代理、科学数据和专家工具。
在Asteromorph的内部评估中,作为基础模型的GLM-5.1在“前沿科学研究”领域得分为10.0%,而应用了推理层的Thila 1.0得分为33.9%。该公司将此结果解读为:不仅基础模型的性能,而且在模型之上构建的推理和验证机制也会影响科学问题解决能力。
该公司还利用去年7月举行的第56届国际物理奥林匹克竞赛(IPhO 2026)的理论考试进行了评估。据该公司称,其自主研发的人工智能系统在不访问外部人工智能模型或互联网的情况下,进行了长达五个小时(与实际参赛者考试时长相同)的推理、验证和答案生成,最终获得了30分中的28.6分。
假设 → 实验 → 数据 → 下一个假设……人工智能研究形成一个“闭环”。
美国公司Lila Sciences正在构建一个“人工智能科学工厂”,该工厂将人工智能模型与基于机器人的自主实验室相结合。它采用闭环系统运行,人工智能生成假设并设计实验,在自动化环境中执行实验,并将结果数据反馈给模型以设计下一个实验。
Layla利用该系统合成了超过95万个mRNA序列,并利用在活细胞中测量的数据训练了一个人工智能模型。随后,该平台被应用于体内CAR-T细胞疗法研究。该公司解释说,在该公司公布的非人灵长类动物实验中,与对照组相比,该平台能够更彻底、更持久地清除B细胞。
Periodic Labs 将人工智能与自主实验室相结合,探索超导体和磁性材料等新型材料。其架构包括利用人工智能分析实验结果,为后续实验提供信息,并利用新生成的数据训练模型。
15日,该公司发布了“周期性氖”(Periodic Neon)模型,该模型拥有1万亿个参数,并基于其自身实验室数据进行训练。该模型专注于通过分析X射线衍射(XRD)数据来识别实验中产生的物质,目前已在实际实验室中部署应用。
探索循环(Discovery Loop)的重点在于利用人工智能(AI)实现研究和工程流程本身的自动化,而非局限于特定的科学领域。它构建了一个人工智能持续处理研究人员重复性任务(例如提出和执行实验、分析结果以及设计后续实验)的架构,并并行执行多个实验以确定下一个研究方向。该计划旨在拓展应用领域,从机器学习的研究和工程入手,逐步扩展到药物开发、医学信息学、太阳能和清洁水等领域。
尽管这些公司的方法各不相同,但它们的共同之处在于,它们都将人工智能从一个回答研究人员问题的工具扩展到一个参与“假设-实验-验证”迭代过程的系统。本质上,人工智能模型推理性能的竞争正在转移到如何有效地将实际实验数据、研究设备和专用工具连接起来的阶段。
Asteromorph 的一位官员表示:“科学人工智能未来的竞争力不仅取决于在基准测试中取得高分,还取决于它如何有效地将提出有意义的假设这一过程扩展到通过实验验证这些假设。”他补充道:“我们将专注于将人工智能推理与实际实验紧密联系起来,从而将复杂的科学研究转化为可重复的系统。”
Asteromorph, une IA qui effectuait des recherches bibliographiques, formule désormais des hypothèses et mène des expériences… Concours « IA scientifique »
Le rôle de l'IA dans la recherche scientifique évolue : d'un outil auxiliaire de recherche et de synthèse d'articles, elle devient un acteur à part entière du processus de recherche. Cette approche implique un raisonnement multidirectionnel pour résoudre des problèmes complexes, la formulation d'hypothèses, puis l'intégration des résultats expérimentaux dans les inférences ultérieures. Dans des domaines tels que la biologie et la science des nouveaux matériaux, où la vérification des hypothèses requiert des expériences concrètes, la capacité à connecter données spécialisées, outils de recherche et environnements expérimentaux devient primordiale, au-delà des performances du modèle d'IA lui-même. C'est dans ce contexte que l'« IA scientifique », qui fait le lien entre l'IA généraliste et la recherche appliquée, émerge comme un nouveau champ de compétition technologique.
Récemment, plusieurs exemples d'utilisation de l'IA pour des processus d'inférence et de vérification à grande échelle ont été observés en mathématiques. Ce mois-ci, OpenAI a proposé une solution au problème du millénaire de Navier-Stokes en déployant environ 10 000 agents simultanés à l'aide d'un modèle de recherche interne. Il a fallu environ 88 heures aux agents pour trouver la solution, suivies de 17 heures supplémentaires pour la formalisation Lean et la vérification à l'aide de GPT-6 Astra.

Antropic a utilisé Claude pour formater la démonstration existante du dernier théorème de Fermat en Lean en 11 jours. Il ne s'agissait pas tant de trouver une nouvelle démonstration que de convertir une démonstration mathématique existante sous une forme vérifiable intégralement par ordinateur. Le résultat final a nécessité environ 13 millions de lignes de code Lean.
En mathématiques, les résultats peuvent être vérifiés par des tests formels, mais la situation est différente en biologie ou en science des matériaux. Même si l'IA propose une hypothèse très probable, l'orientation des recherches futures doit être déterminée par la vérification des résultats au moyen d'expérimentations concrètes et la correction des hypothèses erronées. C'est pourquoi le développement de systèmes reliant les capacités de raisonnement de l'IA aux connaissances spécialisées, aux données, aux outils de recherche et aux processus expérimentaux et de vérification est en cours.
En Corée, Asteromorph développe une « couche de raisonnement » qui relie les modèles d'IA généralistes à la recherche scientifique. Cette structure élabore des hypothèses par raisonnement itératif en combinant des modèles de langage à grande échelle (LLM) avec des données scientifiques spécifiques au domaine, des outils spécialisés et les processus de collaboration et de vérification de plusieurs agents d'IA.
Le système de démonstration « Thyla 1.0 » adopte une structure indépendante de tout modèle de base. Il combine divers agents d'IA, des données scientifiques et des outils experts en fonction du problème à résoudre.
Lors de son évaluation interne, Asteromorph a attribué un score de 10,0 % à GLM-5.1, utilisé comme modèle de base, sur le thème « Frontier Science-Research », tandis que Thila 1.0, avec une couche d'inférence, a obtenu 33,9 %. L'entreprise a interprété ce résultat comme indiquant que la performance en résolution de problèmes scientifiques dépend non seulement des performances du modèle de base, mais aussi de la manière dont l'inférence et la validation sont structurées.
Une évaluation a également été menée à l'aide de l'épreuve théorique de la 56e Olympiade internationale de physique (IPhO 2026), qui s'est tenue en juillet dernier. Selon l'entreprise, son système d'IA propriétaire a effectué le raisonnement, la vérification et la génération des réponses pendant cinq heures – soit la même durée que les participants – sans accéder à des modèles d'IA externes ni à Internet, et a obtenu un score de 28,6 sur 30.
Hypothèse → Expérience → Données → Hypothèse suivante… La recherche en IA devient un « cycle fermé »
Lila Sciences, une entreprise américaine, développe une « usine scientifique d'IA » qui combine des modèles d'IA avec des laboratoires autonomes robotisés. Fonctionnant selon un système en boucle fermée, l'IA génère des hypothèses et conçoit des expériences, les exécute dans un environnement automatisé, puis réinjecte les données obtenues dans le modèle pour concevoir l'expérience suivante.
Layla a utilisé ce système pour synthétiser plus de 950 000 séquences d'ARNm et a exploité des données mesurées sur des cellules vivantes pour entraîner un modèle d'IA. Par la suite, la plateforme a été appliquée à la recherche in vivo sur les cellules CAR-T. L'entreprise a expliqué que, lors d'expériences menées sur des primates non humains et divulguées par ses soins, les résultats ont montré une élimination des lymphocytes B plus profonde et plus durable que celle observée avec le groupe témoin.
Periodic Labs associe l'intelligence artificielle à des laboratoires autonomes pour explorer de nouveaux matériaux, tels que les supraconducteurs et les matériaux magnétiques. Son architecture repose sur l'analyse, par l'IA, des résultats expérimentaux afin d'orienter les expériences ultérieures, et sur l'utilisation des données ainsi générées pour entraîner les modèles.
Le 15, la société a dévoilé « Periodic Neon », un modèle doté d'un billion de paramètres, entraîné sur ses propres données de laboratoire. Ce modèle permet d'identifier des substances générées lors d'expériences grâce à l'analyse de données de diffraction des rayons X (DRX) et a été déployé dans des laboratoires.
Le Discovery Loop vise à automatiser les processus de recherche et d'ingénierie grâce à l'IA, plutôt que de se concentrer sur des domaines scientifiques spécifiques. Il s'agit d'une structure dans laquelle l'IA traite en continu les tâches répétitives des chercheurs — telles que la proposition et la réalisation d'expériences, l'analyse des résultats et la conception d'expériences ultérieures — et mène plusieurs expériences en parallèle afin d'identifier les prochaines pistes de recherche. L'objectif est d'étendre les domaines d'application, en commençant par la recherche et l'ingénierie en apprentissage automatique, à des secteurs comme le développement de médicaments, l'informatique médicale, l'énergie solaire et l'accès à l'eau potable.
Bien que ces entreprises diffèrent dans leurs approches, elles s'accordent sur le fait qu'elles font évoluer l'IA d'un outil répondant aux questions des chercheurs vers un système participant au processus itératif « hypothèse-expérimentation-vérification ». En substance, la compétition en matière de performance d'inférence des modèles d'IA évolue vers une phase où elle se joue sur l'efficacité avec laquelle les données expérimentales réelles, les équipements de recherche et les outils spécialisés peuvent être connectés.
Un responsable d'Asteromorph a déclaré : « La compétitivité future de l'IA scientifique ne dépend pas seulement de l'obtention de scores élevés aux tests de référence, mais aussi de sa capacité à étendre efficacement le processus, de la proposition d'hypothèses pertinentes dans des contextes de recherche réels à leur vérification par l'expérimentation », ajoutant : « Nous nous concentrerons sur un lien étroit entre le raisonnement de l'IA et les expériences concrètes afin de transformer la recherche scientifique complexe en un système reproductible. »