Prospective Students

Our Research: Data Lab works at the intersection of data mining, machine learning, network science, and the social sciences. A common thread in our work is doing more with less: finding reliable, interpretable patterns in data that is large, noisy, incomplete, or unlabeled — and then rigorously verifying that those patterns are real. Our work currently spans three overlapping areas:

  • Misinformation, disinformation, and online behavior. Detecting fake news early and with limited information, modeling how false content propagates, assessing source credibility and the intent of those who spread it, hate speech and hateful memes, and the societal impact of information ecosystems. This line of work is supported by the NSF and appears in venues such as KDD, WWW, CIKM, ICWSM, SDM, and ACM Computing Surveys.
  • Networks and graphs. Interpretable network representations (representing a whole network as its spectral moments, or as a 3D shape), higher-order networks, graph sparsification and graph neural networks, link prediction, network robustness, and noise-enhanced network science — using noise to improve network algorithms rather than degrade them.
  • Machine learning and evaluation. Learning under noise and with limited labels, algorithm stability, interpretable methods, and evaluation without ground truth: how do you know a finding is real when there is no answer key? This is the focus of Reza’s NSF CAREER award.

We often work at scale, and the lab runs its own dedicated computing clusters for training and inference on large language models and graph neural networks.

What We Build: We care about research other people can actually use. A few things that came out of the lab:

  • fake-news.site — a living companion to our fake news survey: tutorials, a research atlas, curated papers, and datasets.
  • datasets.syr.edu — our repository of large-scale social network datasets (19M+ nodes and 169M+ edges across seven kinds of platforms), plus release datasets such as ReCOVery and CHECKED.
  • WebShapes — an interactive tool that turns a network you upload into a 3D shape you can look at (ICDM’18, KDD’20, WSDM’20). Two of the students who built it joined us as summer interns while still in high school, and are co-authors on the paper.
  • Social Media Mining: An Introduction (Cambridge University Press) — our textbook, used in 100+ courses across 30+ countries, and its accompanying tutorial. This is the best place to start if you want to understand how we think about problems.
  • github.com/SUDataLab — our code and data releases.

Our full record is on the publications page. Reading two or three papers close to what interests you is the single best preparation for contacting us.

PhD Students: We are always looking for self-motivated PhD students. Preference is given to students with (1) a solid theoretical and mathematical background, (2) solid programming skills, and/or (3) prior research experience. If the research above is what you would like to spend the next several years on, please apply to Syracuse University’s graduate program in Electrical Engineering and Computer Science (here), mention Reza Zafarani‘s name in your application, and also submit the Data Lab application below so that we can find you.

Undergraduate and MS Students — Volunteer Research Positions: We almost always say yes. If you are an undergraduate or MS student at Syracuse who wants to join the lab as a volunteer research assistant, we accept nearly everyone who is a genuine fit, so please do approach us. Two conditions are firm:

  • You have to fill out the forms. Every request goes through the Data Lab application — choose Undergraduate Research or MS. This is how we track, review, and reply to requests; messages sent any other way tend to get lost. It takes a little effort to complete, and that is deliberate.
  • You have to be a good fit. Fit is about the match between you and a specific piece of our work, not about your GPA.

In practice, a good fit means: you can program (Python; PyTorch, or graph and NLP libraries, are a plus); you have taken or taught yourself the basics of machine learning, data mining, statistics, or network science; you can commit meaningful hours per week for at least a semester, since research on this scale does not fit into two weeks; and you can name a project, paper, or dataset of ours that you want to work on and say why. That last point matters most — tell us what you want to do, and we will tell you honestly whether we have a project for it.

These positions are volunteer, and unpaid, but they are real research. Depending on how things go, they can turn into independent study or research credit, an honors or MS thesis, co-authorship on a paper, and a recommendation letter from someone who has actually seen you work.

Where Our Students Go: Our PhD alumni have gone on to faculty positions (Boise State, UT Austin, DePaul) and to research roles at Amazon and Gallup. Our MS alumni are at Amazon, Google, Microsoft, Meta, Databricks, Visa, and elsewhere, and several have continued to PhD programs, including here at Syracuse. Undergraduates who did research with us have gone on to PhD programs at Cornell, USC, and UC Irvine. Students working with us have won the All University Doctoral Prize twice, along with best paper and best poster recognition at Hypertext and ECS Research Day.

How to Contact Us: Fill out the Data Lab application. Once it is submitted, if we think you are a good fit, we will contact you. General questions that are not about joining the lab can go through our contact page.