ARYAN KARGWAL

[PhD Researcher] AI & Robotics
@ Polytechnique Montréal

Aryan Kargwal

[ THE LAB BENCH ]

Systems Engineering & Research Prototypes

Active Development

BenchYard: Edge Device Benchmarking & LLM Inferencing

Edge devices benchmarking and inferencing for LLMs. Optimizing large language model deployment and performance evaluation across edge computing environments.

LLMEdge ComputingBenchmarking
Research Prototype

Multi-Robot SLAM: Formal Language Verification

Formal language verification approach for multi-agent SLAM systems. Uses temporal logic specifications to ensure correctness guarantees in decentralized mapping and localization tasks.

RoboticsSLAMFormal Verification
Archived

Caiman Ice: Semantic Segmentation for Remote Sensing

SLAM-based semantic segmentation for ice and snow classification in remote sensing imagery. Uses SLAM-derived super-segments merged via visual cues to distinguish between water, snow, and ice in ambiguous conditions.

Semantic SegmentationRemote SensingSLAM

[ SELECTED ACHIEVEMENTS ]

Robotics, AI safety & research recognition

01

Mila · 2026

First Place — Rough-Terrain Quadruped Challenge

Won the Montréal Robotics Summer School challenge by training a quadruped locomotion policy to traverse difficult terrain in an intensive three-day build.

02

CIFAR · 2026

Selected Participant — AI Frontiers School: AI Safety Edition

Selected for CIFAR's interdisciplinary AI safety program for early-stage PhD and advanced master's researchers.

03

Mitacs · 2022

Globalink Research Internship

Completed a funded international research internship at INRS, investigating semantic segmentation for geospatial and Arctic imagery.

04

Mitacs · 2023

Globalink Graduate Fellowship

Awarded a graduate fellowship to continue research in Canada following the Globalink Research Internship.

#01

Agent evaluation metrics: how to measure whether an agent works

Arize AI Aug 2026

A framework for choosing outcome, quality, cost, safety, behavior, and performance metrics that demonstrate whether an agent is actually completing its job.

Agent Ops
p}]]W*>] ;jl(UQ_:jovMO'C\8zo(0Q_
IJ\a^nk'_mYvnpjjqh!oQox.8[Y|qipa
+{I?mW~czO<c@#_@J]xB>_vI"n`"M?@~
X(-d YYjU[\jwof\>@BcY`pa|(vwoo''
^_0wn)|jvlxZ,x ?^W1UQ~M|qZ-[aatJ
M*!U~l>U`)]_)^(;&^/tZl_p}:qW1/1~
+;ah@a'||lB{wU]]{W M^w@B`-O^ L/^
[fa|j}%m<qpkJ$]QM|CC/[&%o`Mr0)\0

BETA DISTRIBUTION

#02

Agent observability: how to trace, debug, and improve AI agents

Arize AI Aug 2026

A practical guide to capturing agent traces, connecting execution across sessions, and using evaluations to diagnose and improve production AI agents.

Agent Ops
                    {${${${${   
    { {      {  {{ {$Q$Q$Q$Q${{ 
 { {${${ {  {${{$${$QQQQQQQQQ$${
{${$Q$Q${${{$Q$$QQ$QQQQQQQQQQQQ$
$Q$QQQQQ$Q$$QQQQQQQQQQQQQQQQQQQQ
QQQQQQQQQQQQQQQQQQQQQQQQQQQQQQQQ
QQQQQQQQQQQQQQQQQQQQQQQQQQQQQQQQ
{{{{{{{{{{{{{{{{{{{{{{{{{{{{{{{{

MARKOV CHAIN

#03

How to evaluate AI agents: a production workflow

Arize AI Aug 2026

A production workflow for evaluating agent outcomes, trajectories, individual decisions, and reliability, then turning failures into regression tests.

Agent Ops
    >>>>>>>   >    > > >>>>    >
   >   >>  > > > >> > >    >  > 
  >   >  >  >   >>>         >>  
 Y > >    >     >  > > >    >|  
$ b >      >> >> |  > > |>>>> | 
 | Y | > >>>Y>|>|>|>   >||||>  Y
    Y Y >|>| Y |> >|>>>| >>>>>>>
     > || >   >     |||      >> 

BROWNIAN MOTION

#04

AI agent frameworks compared: LangGraph, CrewAI, AutoGen, and more

Arize AI Aug 2026

A practical comparison of agent SDKs, orchestration runtimes, harnesses, and managed platforms for production AI systems.

Architecture
                  >> > >  > >>  
   |             >  >>> >> >> > 
  | | Y > | >   |   > > >> >>> >
 b   Y Y Y | > Y > > > |  >> >>>
$ | >   | > | Y>|>Y > > > |   |>
 > |    > >>>Y>>>|>| > > | > > |
  > |  >>> >>>>    >> > |   > > 
   > ||> >>>        >>>> >>>>>  

BROWNIAN MOTION

#05

Best AI Observability Tools for Autonomous Agents in 2026

Arize AI Feb 2026

A comparative look at leading observability tools for autonomous AI agents, with evaluation criteria for production deployment in 2026.

Tools & Platforms
?i!i">>i+_+{+!-[+_+':+"><!,`.><!
i;l>l,>li!l>I]!I:>~>iI<`:>>+<:i`
.->-,;^:!_}[f|)\xX|(~1)<}<,><"~!
?,+-l??+}\njzOCJLqpUcf|[1~[_<i I
!_><{!}/Xruj#*aB$kZZ0zUtv/[?}i><
l?>]-?<]}jr|QLzqY0Czr|Y1?[?!>!?<
^_I~~+1:]+1]rjr|ff~\<[~|+-<>~i+?
> i>>{,il[~<!+-+?<1?<l+]~:~l><,l

GAUSSIAN DISTRIBUTION

#06

Understanding Latency in AI Model Deployment

Pipeshift Feb 2026

A practical breakdown of where latency comes from in deployed AI systems and how teams can reduce bottlenecks across the inference path.

Engineering
 ^ ' '.`' '''`''` .'`'. '`'`  ''
-+ii;,;^`"^`.^^ .^' `^  .`'`` '^
v/[]>!;l;"`'^"^'`^.`^`.'.` `'.'`
bCn\]-<>:;"""`^^`..``.. `'^.^'^'
$kUx\[?_li;"^,`'```'^^' ^..  . .
wCr\]]+<Il:;`^'''^'.`' .'``  ^  
n|{+~i;I";^`"`^`^..'`'  `   ` ` 
_+>I:I;^""^.`` .. .^.``'^`^ .^  

EXPONENTIAL DECAY

[ Currently Reading ]

View on Goodreads →