Shahriar Shayesteh

PhD Candidate in Informatics · Pennsylvania State University

Seeking Summer 2027 research and applied-science internships.

I am a PhD candidate in Informatics at Penn State, advised by Shomir Wilson in the Human Language Technologies Lab.

My research focuses on two complementary areas: agentic AI privacy—how LLMs, tools, and context shape privacy risks when agentic systems act on users’ behalf—and NLP for consumer-oriented legal documents—how language technologies can help consumers navigate contracts and policies, and the failure modes that can limit their usefulness.

Shahriar Shayesteh
Human Language Technologies Lab
Penn State

Research

AI systems increasingly mediate information between people, digital services, and complex documents. I study this role through two NLP-centered research areas: privacy in agentic AI systems and language technologies for consumer-oriented legal documents.

Agentic AI Privacy

Agentic AI systems combine LLMs, tools, context, and system orchestration to act on users' behalf. I study how these components shape personal-information disclosure, data minimization, and privacy risks when agents communicate with third-party services.

Tool-Using LLMs · Data Minimization · Contextual Integrity · Information Disclosure
Diagnostic audit of 2,344 tool schemas · Controlled behavioral evaluations ongoing

News

Privacy, More or Less is under revision for PoPETs 2027.

Invited to give a guest lecture in IST 574: Human Language Technologies at Penn State.

Presented From Conventional Web Privacy to Agentic Disclosure at PrivateNLP @ ACL 2026.

SoACer published at ACM DocEng 2025.

PrivaSeer presented at USENIX SOUPS 2025.

Selected Publications

From Conventional Web Privacy to Agentic Disclosure: How Tool Schemas May Invite LLM Oversharing

Shahriar Shayesteh, Shomir Wilson

PrivateNLP @ ACL 2026

SoAC and SoACer: A Sector-Based Corpus and LLM-Based Framework for Sectoral Website Classification

Shahriar Shayesteh, Mukund Srinath, Lee Matheson, Lu Xian, Sinjoy Saha, C. Lee Giles, Shomir Wilson

ACM DocEng 2025 · Corpus of 195,495 websites

The PrivaSeer Project: Large-Scale Resources for Analysis of Privacy Policy Text

Shomir Wilson, Florian Schaub, Lee Matheson, Shahriar Shayesteh, Lu Xian

USENIX SOUPS 2025 · Poster · 3,967,487 policies as of April 2025

Generative Adversarial Learning with Negative Data Augmentation for Semi-Supervised Text Classification

Shahriar Shayesteh, Diana Inkpen

FLAIRS-35 · 2022

See full publication list on Google Scholar →

Selected Research Projects

How Tool Schemas Shape LLM Agent Disclosure: Structure, Guidance, and Declared Purpose

Ongoing research

Uses controlled schema interventions to test effects on unnecessary personal-information disclosure while holding the task and backend operation fixed.

Easier to Read, Still Easy to Miss: Auditing Content Allocation in LLM Explanations of Wireless Contracts

Manuscript in preparation · Target: ACM FAccT 2027

Uses 24 topics and 3,899 verified provisions to evaluate allocation, omission, scope preservation, and factual support.

Privacy, More or Less: A Large-Scale Comparison of Privacy-Policy Disclosure Patterns Across Online Sectors

Under revision for PoPETs 2027

Analyzes 59,590 matched 2022–2025 policy pairs for cross-sector changes in recipient specificity, purpose coverage, and disclosure opacity.

Talks & Teaching

Invited guest lecture (scheduled), IST 574: Human Language Technologies, Penn State.

Presented From Conventional Web Privacy to Agentic Disclosure at PrivateNLP @ ACL 2026, San Diego.

Participant, Mila Responsible AI and Human Rights Summer School, Montréal.