Research

Publications, labs, and research communications.

Research Focus

My research sits at the intersection of intelligent systems, algorithms, and applied machine learning. I build computational tools and experimental frameworks for privacy-preserving ML, distributed systems simulation, and biological data modeling, while also contributing to climate-focused evidence synthesis and reproducibility analysis. Across projects, I prioritize rigorous evaluation design, reproducible pipelines, and systems-aware implementations that can scale beyond toy settings.


Publications

Selected publications and preprints. I include DOI, preprint, and PDF links where available.

A Systematic Literature Review of Climate Econometrics Research

Preprint (SSRN) ()

PDF • Preprint • DOI


Current and Recent Research Work

Selected systems, simulations, and algorithmic frameworks developed through lab collaborations.

Simulated Annealing Haplotype Aligner (SAHap)

A C++ haplotype assembly tool that uses simulated annealing and super-read strategies to reduce weighted mismatch error and improve reconstruction quality.

Dr. Wayne Hayes Research Group @ UCI • Sep 2024 - Sep 2025

Led the simulated annealing-based haplotype assembly project as a freshman researcher, implemented local optimization strategies (including simulated annealing and hill climbing), and presented the work at the UCI Undergraduate Research Symposium.

Related link


Privacy-Preserving Deletion Benchmarking Pipeline

Python and MySQL-backed experimental pipeline for evaluating differential privacy and ILP-based deletion mechanisms under Right-to-Be-Forgotten settings across UCI ML datasets.

Mehrotra Lab (Sharad's Lab) @ UCI • Aug 2025 - Present

Designed and implemented deletion mechanisms, evaluated RTBF guarantees through controlled studies, and built a reproducible benchmarking framework for deletion efficiency and leakage analysis in collaboration toward a VLDB submission.


Large-Scale Sharding Simulation Framework

A discrete-event simulator for vector-search sharding strategies (including IVF and ANN variants) used to study bursty workloads, failure propagation, and graceful degradation.

Network, Systems, and AI Lab (NetSAIL) @ UCI • Aug 2025 - Present

Built simulation infrastructure to analyze system resilience, load imbalance, and degradation behavior under server failures and non-stationary workloads.


Biological Model Evaluation and Soft-Verifier Framework

Graph-metric evaluation pipelines and a soft-verifier training framework for weakly supervised biological ML/LLM systems over gene relationship datasets.

Zhang Lab @ UCI • Aug 2025 - Present

Investigated gene relationship datasets used in biological ML/LLM models, designed graph-based evaluation metrics across cell-line datasets, and developed a soft-verifier framework to reduce dependence on proprietary data.


Research Presentation

Talk or overview video with abstract.

Simulated Annealing Haplotype Aligner

A faster and novel approach to genome haplotype assembly using simulated annealing and super reads.

Presenter(s): Arnav Dhariya, Pranavi Gollanapalli

Faculty Mentor(s): Wayne Hayes

Abstract

Simulated Annealing Haplotype Aligner (SAHap), a C++ application, tackles the challenge of separating parental genome sequences from lossy end reads. Rather than merging all reads into a single genome, similar to most genome sequencing algorithms, it treats each read as belonging to one of two haplotypes (or more, in polyploid cases). The system seeks the configuration that minimizes the total weighted mismatch between reads and their assigned sequence (wMEC).

Reads are already sequenced into the correct positions (sites). SAHap shuffles haplotype assignments randomly and then assesses the resulting change in error. It gradually tightens the search: swaps that reduce overall wMEC are always kept, while swaps that increase error can still be accepted early in the run to avoid stagnation in local minima, similar to issues seen in hill-climbing approaches. A temperature schedule gradually phases out acceptance of worse moves, guiding the program toward a stronger arrangement of reads into haplotype groups.

At present, the algorithm minimizes error successfully but can leave high-wMEC segments that do not reach the optimal solution. To smooth errors in these regions, SAHap fuses consecutive, vertical, perfectly matching reads into longer super-reads (based on the MaSURca assembler), reducing the number of variables and enabling better solutions. Another proposed improvement is rerunning simulated annealing on windows with high wMEC. When tested on datasets of five hundred and then twenty-thousand reads, this two-stage approach with super-reads is expected to deliver lower error scores while improving runtime.

Labs

Groups and labs I’m affiliated with.

  • Mehrotra Lab (Sharad's Lab) @ UCI — Undergraduate Researcher (Aug 2025 - Present)
  • Zhang Lab @ UCI — Undergraduate Researcher (Aug 2025 - Present)
  • Network, Systems, and AI Lab (NetSAIL) @ UCI — Undergraduate Researcher (Aug 2025 - Present)
  • Dr. Wayne Hayes Research Group @ UCI — Research Team Lead / Undergraduate Researcher (Sep 2024 - Sep 2025)
  • Green IT Lab — Literature Reviewer / Researcher (Jan 2025 - Mar 2025)
  • Computer Vision Lab @ UCI — Undergraduate Researcher (Aug 2025 - Jan 2026)
  • INsite Lab @ UCI — Research Team Lead (Jun 2025 - Dec 2025)