Yixuan Zhang
Champaign

Open to Summer 2027 · SWE / MLE internships

Yixuan Zhang builds ML systems that hold up outside the notebook. Private split inference, edge vision and LLM serving.

nowMaster of Science in Computer Science · UIUC ’28beforeB.Sc. Computer Science · Nottingham, First Class

σ 0.31 shape 16×16×16
FIG. 0 The client-side tensor from my thesis, with Gaussian noise before it leaves the device. Move across it to change σ.
Top-1 · FaceScrub
81.96%

+1.63 pts over the published CEM defence

Jetson Orin Nano
220→95ms

2.3× faster, 6.2× smaller, <0.5 mAP loss

Agent context
−85.5%

tokens per query at 0.984 top-1 routing

B.Sc. Nottingham
3.9/4.00

First Class, ranked 1st in CV and Security

Selected work 10

Click a row for details

01 DualPath-CEMResearchSplit inference that resists model inversion without giving up accuracy. 81.96%top-1 · +75.8% attack MSE PyTorch2025–26
  • Two client paths: a Gaussian-protected 16×16×16 SlotCEM tensor carries the privacy burden and a 256-D MobileNetV3 token carries utility. The server fuses them with calibrated logits.
  • On a 44,139-image FaceScrub split (526 identities, 5 seeds), top-1 rose 1.63 points over the published Noise_ARL+CEM reference while decoder and GAN reconstruction MSE rose 75.8% and 74.9%.
  • Manuscript in preparation for Pattern Recognition with Prof. Jianfeng Ren. Full write-up in Research.
github.com/asher0913/DualPathCEM ↗
02 SkillRouterLLM agentsProgressive tool disclosure for large skill catalogues. −85.5%tokens · top-1 0.984 PythonAug–Sep 2026
  • Three-tier registry: routing reads only metadata, and full instructions load only for the skills actually selected.
  • IDF-weighted lexical prefilter protects recall, then a ranker selects top-k. On a 60-skill catalogue, context fell from 14,108 to 2,052 tokens per query at top-1 0.984 and top-3 1.000.
  • Also measured on held-out paraphrases that avoid every trigger phrase: lexical routing drops to 0.21 top-1, and an optional MiniLM stage tuned only on the dev set brings it to 0.56 while declining 85% of out-of-domain questions.
  • Content fingerprints and a CI admission gate block catalogue changes that add trigger conflicts or drop routing accuracy. 30 tests in GitHub Actions.
github.com/asher0913/skill-router ↗
03 InferenceLabLLM servingDynamic batching, request coalescing and a semantic cache in front of vLLM, SGLang or Ollama. 16→3backend generations asyncio · vLLMSep–Oct 2024
  • Async dynamic batcher with batch-size and queue-wait bounds, single-flight coalescing of identical in-flight prompts, and a TTL/LRU semantic cache with a guard on numbers, quoted text and negation.
  • On an 80-request, 16-concurrency workload, coalescing cut backend generations from 16 to 3 at an 80% cache-hit rate on the deterministic backend.
  • A labelled cache study found semantic caching serving wrong answers, such as swapped unit conversions; per-family thresholds answer 67.6% of paraphrases with none wrong. A scheduling simulation shows continuous batching sustaining 10× the load of static batching within a 10 s p95 SLO.
  • Swappable OpenAI-compatible backend, so the same harness runs against real vLLM, SGLang or Ollama servers.
github.com/asher0913/inference-lab ↗
04 SecureRAGRetrieval securityMulti-tenant hybrid retrieval that checks permissions before ranking. 54%→0%leaks after ACL changes FastAPI · RAGOct–Nov 2023
  • Tenant and ACL filters run before hybrid ranking, with collection statistics from the caller's view only; each candidate is then re-checked live against the directory of record, failing closed, so revoked text never reaches the model.
  • Benchmarked five enforcement designs on a two-tenant corpus: post-filtering left 25% of contexts empty, and after permission changes a pre-filter on the stale index leaked on 54% of queries against 0% with the live re-check.
  • Citation and denial audit events on every query.
github.com/asher0913/secure-rag ↗
05 Multi-Material Resin SlicerC++ desktopSTEP/STL assemblies to per-material masks and G-code. C++17Qt · OpenGL · headless CI test InternshipMay–Aug 2026
  • Imports STL and STEP assemblies through OpenCascade while preserving nested transforms, with per-part material assignment and an OpenGL VBO preview.
  • Rasterises per-material PNG masks in a background worker with cancellation and timeouts, plans tank changes layer by layer, and writes merged exposures plus G-code.
  • A headless --selftest runs the full import-to-G-code path under Xvfb in Ubuntu CI. Packaged for macOS and Windows.
github.com/asher0913/multi-material-slicer ↗
06 AgentGuardAgent safetyA policy gateway between an LLM and the tools it wants to call. 100%precision / recall on fixture FastAPIJan–Feb 2025
  • Every proposed tool call returns allow, require-approval or deny, with a bounded risk score, readable reasons and stable rule IDs for audit.
  • Covers destructive shell commands, broad filesystem targets, credential exfiltration, prompt-injection propagation and multi-step escalation.
  • 100% precision and recall on the checked-in adversarial fixture.
github.com/asher0913/agent-guard ↗
07 CodeTaskForgeAgent evaluationThe execution and evidence layer for coding agents. 25→0wrong patches accepted Python · DockerSep–Oct 2025
  • SWE-bench-style judging: a path policy, throwaway workspaces, shell-free runs with scrubbed environments and process-group timeouts, restored tests, fail-to-pass and pass-to-pass, and held-out tests.
  • On 98 labelled candidate patches, exit codes accepted 25 wrong ones (test tampering, hard-coded answers), the SWE-bench protocol 15, and held-out tests none; every rejection was classified into the right one of eight failure kinds.
  • CLI, FastAPI and Docker share one harness; the API never reveals hidden test names.
github.com/asher0913/code-task-forge ↗
08 Contactless Breathing MonitoriOS · on-deviceTrueDepth and Vision pose landmarks turn chest depth into a live breathing curve. 0 bytesleave the device SwiftMar–Apr 2024
  • AVCaptureMultiCamSession synchronises front-camera video with TrueDepth depth, and Vision shoulder landmarks define the chest region of interest.
  • Gaussian filtering and a weighted moving average turn the cropped depth signal into a respiratory curve plotted live in SwiftUI, and an autocorrelation estimator reports breaths per minute (within 1 breath/min on 99.6% of noisy synthetic signals).
  • No network client and no persistence. Grew out of a Software Engineering group project I led under Prof. Matthew Pike.
github.com/asher0913/contactless-breathing-monitor ↗
09 Oxford-IIIT PetFine-grained vision37-breed classification with transfer learning, EMA, TTA and a hand-written Grad-CAM. 89.8%test accuracy, ResNet-18 PyTorchNov–Dec 2024
  • Fine-tuned ResNet-18 with mixed precision, cosine scheduling, EMA and horizontal-flip TTA to 89.8% test accuracy on 37 classes.
  • A from-scratch SE-ResNet reached 49.7% as the controlled baseline, with ablations, confusion matrices and Grad-CAM overlays.
github.com/asher0913/oxford-pet-classification ↗
10 IAMABOTMechMania 32A 24-hour AI contest bot built from the game engine’s Rust source. 7th→3rdon the live leaderboard PythonSep 2026
  • Commands a fleet of up to 32 bots on a walled 32×32 map. Every rule the bot relies on was read from the engine’s Rust source rather than the prose rules.
  • Engine-exact fire control, payload control and a replay-driven regression arena for testing strategy changes.
github.com/asher0913/IAMABOT ↗

Labs and coursework 20

Laptop-scale reference implementations of production problems, built for reproducible evaluation.

tiered-kv-cache-labHBM → DRAM → NVMe prefix caching, TTFT under load, SLO capacity planningserving ai-gateway-control-planeBudgets, circuit breakers, latency-aware routing: p99 14.3 s → 5.8 s under faultsserving edge-quantization-labINT8/INT4 calibration, 6.7× searched mixed precision and an integer-only kerneledge vision-pipeline-orchestratorBackpressure-aware CPU/GPU DAG scheduler with retries and a DLQsystems agentopsIncident investigation under flaky telemetry: 0.4% wrong vs 33.5%, replayable tracesagents agent-trace-labAnti-pattern detection and root-cause attribution, stress-tested on 1,000 tracesagents reflective-agent-labSandbox verification cuts irreversible side effects from 12.0 to 0.2 per 100 tasksagents graph-sop-agentService graph plus runbooks vs document-only RAG, with review-gated ingestionagents risk-agent-workbenchEvidence-grounded risk review that abstains when evidence is missingagents document-vlm-data-engineFigure matching, table validation, cross-page QA and hard samplesdata multimodal-safety-labCross-modal red-team benchmark with safety–utility metricssafety recsys-ranking-labALS retrieval and GBDT re-ranking on MovieLens-100K, +40% HR@10ranking plant-leaf-recognitionResNet-101 + ViT-B/16 fusion, 98.46% reported top-1 on 100 leaf speciesvision music-emotion-regressionDEAM valence and arousal across eight model families, R² 0.604ml cifar10-ml-benchmarkPCA, an MLP and a Random Forest under matched CV; PCA helps one, hurts the otherml classical-ai-searchSix search algorithms on 300 seeded mazes, turn-aware routing, MDPsalgorithms bin-packing-metaheuristicsBFD, Minimum Bin Slack and annealing on Falkenauer's OR-Library instancesalgorithms javafx-platformerJava 21 game built on MVC, commands, state machines and data-checked levelssoftware treasure-hunt-pathfindingA* and BFS hints in a Swing game; A* expands 7× fewer cellsalgorithms flask-music-libraryFlask and SQLite catalogue with CSRF-protected, transactional writessoftware

Research

Privacy in collaborative inference

Pattern Recognition Manuscript in preparation · 2026

DDP-CEM: Dual-Path Privacy Protection for Collaborative Face Inference

Yixuan Zhang, Jianfeng Ren

A split model sends an intermediate representation from a trusted client to an untrusted server. Strong noise makes that tensor hard to invert but also removes what the server needs for recognition. DualPath-CEM separates the two jobs: a Gaussian-protected spatial path carries the privacy burden, a compact semantic token carries utility, and the server fuses two classifiers with calibrated logits. The attacker is assumed to see both released tensors.

Top-1
81.96%
Δ vs CEM
+1.63 pts
Decoder MSE
+75.8%
GAN MSE
+74.9%
DualPath-CEM architecture A private image feeds two client encoders. The SlotCEM client produces a 16 by 16 by 16 tensor with Gaussian noise sigma 0.31; the MobileNetV3-Large semantic client produces a 256-dimensional token with noise sigma 0.10. Both cross the trust boundary to an untrusted server where a spatial head and a token classifier are fused with calibrated logits into an identity prediction. An inversion attacker on the server observes both tensors. CLIENT · TRUSTED SERVER · UNTRUSTED Private imageFaceScrub 64×64 SlotCEM clientVGG11-BN · 16×16×16 Semantic clientMobileNetV3-L · 256-D σ = 0.31 σ = 0.10 Spatial serverclassifier head Inversion attackersees both tensors Token classifier256-D → logits Calibrated fusion→ identity
FIG. 1 Two released tensors, one untrusted server. SlotCEM shapes the spatial encoder during training only; no prototypes are transmitted.
TABLE 1 FaceScrub, 526 identities, 44,139 images, mean of 5 seeds. Higher reconstruction MSE means a worse attack.
MethodTop-1 ↑Decoder MSE ↑ inferenceGAN MSE ↑ inference
Noise_ARL + CEM published 80.33% 0.0211 0.0231
DualPath-CEM this work 81.96% 0.0371 0.0404
Evidence boundary

The reference row is the published CEM result, not a local reproduction. All five adaptations start from the same frozen SlotCEM foundation, so they are not independent end-to-end runs. The completed snapshot is FaceScrub-only. The evidence package covers 70 main attack runs, 15 repeated utility evaluations, 96 supplementary attack runs, a matched-capacity control and a ResNet-18 control.

B.Sc. dissertation First Class · Oct 2025 – Jun 2026 · Advisor Prof. Jianfeng Ren

SlotCEM: slot-attention conditional entropy regularisation against model inversion

Replaces the Gaussian-mixture estimator in CEM with a training-only, class-conditioned Slot Attention module. 4,096-D smashed features are projected to 64-D, grouped by class with a bounded per-class memory bank, and eight slots model multimodal structure under a soft geometric-variance loss. Structured channel pruning keeps the edge-side VGG11-BN client light.

TABLE 2 FaceScrub, 530 identities at 64×64, cut layer 4, 16-channel bottleneck. MSE rises 29.8% for a 0.84-point accuracy cost; the dual-path design removes that cost.
MethodTop-1 ↑MIA MSE ↑SSIM ↓PSNR ↓
No defence sanity89.43%0.000670.94931.76
Noise_ARL + CEM published80.33%0.0211——
SlotCEM79.49%0.02740.51815.63

Experience

4 internships · 2 research assistantships

  1. 2026May – Aug

    Ningbo AST Machinery Technology Ningbo

    Software Engineer Intern

    • Shipped a cross-platform multi-material resin slicer in C++17, Qt and OpenGL, with CMake builds and a headless Xvfb end-to-end self-test.
    • Added GPU render caches, shared mesh storage, cancellable background export and malformed-geometry validation. Packaged macOS and Windows builds.
    • Designed a 51-table MySQL CRM with 107 foreign keys and 70 constraints, and shipped 10 Vue/FastAPI modules with role-based access and concurrency-safe lead assignment.
    • Built an LLM research pipeline that produced 345 faculty profiles and 1,127 deduplicated grant records.
  2. 2025May – Aug

    Zhiyang Innovation Technology Shanghai

    Machine Learning Engineer Intern

    • Cut Jetson Orin Nano latency 220 → 95 ms and model size 87 → 14 MB at under 0.5 mAP loss, exporting to ONNX and serving TensorRT engines behind FastAPI and Docker.
    • Raised mAP@0.5 0.74 → 0.79 and recall +6.2 pts while cutting false positives 12% on a 7:1 imbalanced defect set, benchmarking YOLOv8s, YOLOv11, Faster R-CNN and EfficientNet-B3 on 2×A100.
    • Built the training pipeline for 120K+ images from three sources. Automated label QA and augmentation cut label noise 18%, data prep from 3 days to under 1, and added 5.6 mAP@0.5.
    • Clustered false positives with K-means to surface long-tail failure modes from daily inference logs for retraining.
  3. 2024–25Sep – Jul

    University of Nottingham, Sport Department Nottingham

    Data Analyst Intern

    • Wrote Python workflows that clean and aggregate event participation and feedback data, surfacing participation, satisfaction and engagement trends.
    • Built an Excel work-hour system with formulas and automation scripts that tracks attendance, flags incomplete records and generates standard reports.
  4. 2024May – Jul

    Beijing Corefire Technology Beijing

    Software Engineer Intern

    • Built merchant onboarding, query and reporting services for UMFintech’s acquiring platform in Spring Cloud, Spring Boot and MySQL, covering POS and QR transactions, terminal lifecycle and profit sharing.
    • Implemented phone/password and SMS-code login with Redis-backed session tokens, plus self-service merchant registration.
  5. 2025Jan – Jun

    Research Assistant, medical image analysis Nottingham

    Advisor Dr. Kian Ming Lim

    • Trained ResNet and Vision Transformer classifiers on MedMNIST with OpenCV augmentation, then distilled a lightweight student model.
  6. 2024Feb – Apr

    Research Assistant, news sentiment for trading Nottingham

    Advisor Dr. Tianxiang Cui

    • Aligned 100K+ articles from 20+ sources, fine-tuned BERT-Large into a daily sentiment index, and backtested a PPO trading agent on the CSI 300.

Education

And honours

2026 – 2028

University of Illinois Urbana-Champaign

Master of Science in Computer Science

Machine Learning Systems, Distributed Systems, Advanced Data Management, Advanced Topics in NLP, Advanced Operating Systems, Advanced Computer Security.

2022 – 2026

University of Nottingham Nottingham

B.Sc. (Hons) Computer Science · First Class · GPA 3.9 / 4.00

Ranked 1st in Computer Vision and in Computer Security. 4.0 in Operating Systems and Concurrency, Algorithms and Data Structures, Machine Learning, AI Methods, Databases, Systems and Architecture. Advisors Prof. Jianfeng Ren and Prof. Matthew Pike.

2024

Aarhus University Denmark

Summer school · Software Design Using C++

Taught by Bjarne Stroustrup, the creator of C++.

Honours 10

  • Outstanding Graduate of Zhejiang Province2026
  • Outstanding Graduate, top 5%2026
  • Dean’s Scholarship, top 10%2025
  • Outstanding Student, top 5%2025
  • Zhejiang Provincial Government Scholarship, top 1%2024
  • MCM Honorable Mention2024
  • MathorCup Big Data Challenge, Second Prize2023
  • National English Competition, Second Prize2023
  • Nottingham Advantage Award2023
  • Xinghuo Freshman Debate Tournament, Champion2023

About

Stack, languages, and life outside code

Languages
Python, C++, Java, SQL, TypeScript, JavaScript, Go, Swift
ML
PyTorch, Hugging Face Transformers, scikit-learn, OpenCV, YOLO, ViT, LoRA, RAG
Inference
ONNX, TensorRT, vLLM, Jetson, INT8 quantisation, batching, caching
Systems
FastAPI, Spring Boot, Docker, Kubernetes, MySQL, Redis, MLflow, GitHub Actions, AWS, GCP, Linux
Desktop · iOS
Qt, OpenGL, CMake, SwiftUI, AVFoundation, Vision
Spoken
Chinese, native · English, IELTS 7.5 · Japanese, JLPT N3

Leadership

President, Nottingham Debate Union 2023–24Designed a Nottingham Advantage Award course on British Parliamentary debate and led 5 instructors teaching 200+ students.

Deputy Director, Students’ Union 2023–24Ran a 400-person cultural season and career sessions with KPMG, PwC, ByteDance, NIO, Xiaomi and Lufthansa.

Peer Mentor, Computer Science 2023–24Taught 20+ first-years Java, data structures and Git. 90%+ rated it helpful.

Service and campus

Peer Listener, Mental Health Centre 2024Supported 13 students through peer listening.

Chinese Teaching Volunteer, Confucius Institute 2022–23Weekly one-to-one lessons for international students.

Young Volunteers Association · Media Centre · Chinese Debate Society · Chinese Corner 2022–23Recruited 80+ volunteers, wrote for the official WeChat account, and competed in debate.

Working on ML systems, inference or infrastructure? I’d like to hear about it.