I look for what breaks before a system is trusted.

I’m Vicky Feliren, an applied scientist working where AI safety, uncertainty, and underrepresented data meet. My work asks a practical question: when a model becomes safer, does it remain useful for the people and inputs its training represented least?

Explore the research View curriculum vitae
Vicky Feliren
Applied Scientist · AI safety and reliable multimodal systems
7 Peer-reviewed papers Published across language and vision venues, with methods and results reviewed by other researchers.
5+ years Industry ML in production Models used for finance, safety, and city services, where mistakes had real costs.
1 Issued patent A method that passed formal examination and became an issued patent.
12 Awards & scholarships Including the Asia-Pacific regional award at the Global South AI Safety Hackathon.
100+ Researchers co-led with Experience coordinating research across a large international group.
1000+ Students taught & judged Experience explaining difficult ideas to people seeing them for the first time.

I'm an applied scientist studying calibration under safety alignment. Safety training can change how well a model's confidence matches its accuracy. Poorly calibrated models also defer poorly, which limits the tasks we can trust them with. I want safe models to stay useful. My starting position is that the current trade-off can be improved.

At Monash University, I study conformal prediction for vision-language navigation. I work with Associate Professor Risqi Saputra and Professor Taufiq Asyhari. The work taught me to ask what each guarantee depends on and what might break first. I now want to apply that tool here. Even when a model's calibration slips, its decision to defer should have a clear bound.

I also bring experience from Southeast Asia, one of the world's most varied regions in language and visual data. The region is still largely missing from the datasets and benchmarks used to judge modern AI. I have spent years building local datasets with SEACrowd. My research also develops regional adaptation methods. Most safety training data is English text. If its costs vary across languages and inputs, Southeast Asia is a good place to look. Measuring them requires data and native fluency. I have both.

01

The question

A safe model should know when its answer can be trusted.

Safety training can change how well a model’s confidence matches its accuracy. Once that calibration slips, the model may continue when it should defer—or abstain so often that it is no longer useful. Average benchmark scores can hide where that trade-off is being paid.

I study the distribution beneath the average: which languages, input types, and communities absorb the largest cost. I am especially interested in Southeast Asia, where the world’s linguistic and visual variety is still poorly represented in mainstream datasets and evaluations.

02

The path

Production systems taught me what a research guarantee must survive.

Jakarta Smart City

Forecasting municipal waste taught me that model quality matters only when it changes a real allocation decision.

Finance and identity

Biometrics, credit, and fraud systems made calibration, auditability, and failure costs operational—not theoretical.

Earth observation

Satellite systems across sensors and regions made distribution shift visible in every map.

Southeast Asian AI

SEACrowd connected the technical problem to the missing languages, cultures, and visual worlds behind it.

Calibration under alignment

I bring those threads together: measure the hidden cost, then recover useful deference with guarantees.

The complete chronology lives in the CV
03

The method

Start with the failure boundary, then build back toward use.

01

Find the hidden average

Disaggregate the result until the users and inputs carrying the cost become visible.

02

Make uncertainty legible

Turn confidence into a measurable decision variable—not a decorative score.

03

Test outside the comfortable case

Use multilingual, multicultural, and multimodal inputs that expose brittle assumptions.

04

Recover with a bound

Prefer interventions whose limits can be stated clearly enough for someone else to trust.

04

The direction

Can aligned models keep calibrated judgment beyond English and beyond text?

Current research direction

I am measuring how alignment changes calibration across languages and modalities, then testing whether distribution-free abstention can recover reliable deference without erasing usefulness.

Measure
Calibration tax by input group
Intervene
Bounded abstention after alignment
Evaluate
Multilingual + multimodal systems

The work is public. The complete record is separate. Choose the depth you need.

This Month · updated July 2026

  • Writing my M.Sc. thesis on conformal prediction for vision-language navigation, with a focus on abstention and set efficiency
  • Reproducing a published result on the calibration cost of safety alignment in a setup I can run from start to finish
  • Designing the next experiment to measure that cost by language, starting with one I speak natively so I can review the evaluation data myself
  • Presented my multilingual VLM abstention study at AI Safety India's Hackathon Winners event

Research Agenda

Making a model safer has a cost. Researchers call it the alignment tax. Studies usually report one average across a test set. I think the cost varies across inputs, and we have not measured that pattern. It may be highest where safety data is scarce. If so, a model can appear safely aligned while working unevenly across languages and input types. Our current tests may miss that gap because they focus on well-represented data.

Calibration is where the safety tax gets paid

Safety training changes a model's answers, including its confidence in them. Its effect on capability is well studied. Examples include Lin et al. (EMNLP 2024) on the alignment tax and Huang et al. on reasoning. Its effect on calibration has received less attention. Leng et al. (ICLR 2025) find that RLHF makes models express too much confidence. Hu et al. (ACL 2026 Findings) describe a severe loss of calibration. I care about this cost because calibration shapes how much work we can safely delegate.

The cost is probably not spread evenly

Most safety training data is English text. Its costs may vary across inputs that appear in the data at very different rates. I therefore measure calibration changes by language and input type. One average cannot describe a distribution that the test barely sampled. I am running this experiment now, and either result will help. A uniform cost would make the problem simpler. An uneven cost would show that some users receive a less calibrated model than standard benchmarks suggest.

Getting the deference back, with something you can bound

Measuring the problem is only part of the work. We also need a way to recover. Hu et al. (ACL 2026 Findings) recover some calibration by merging model weights from before and after alignment. I want to test an abstention layer after calibration has already declined. The layer would use distribution-free coverage (ICML 2024). I also want to measure the cost to usefulness. Here, uncertainty quantification is a practical tool for a specific failure.

Where it gets tested

These ideas must also work on systems that were never designed for the test. I begin by reproducing a published result in a setup I can run from start to finish. That gives me a sound result to build on. I then use inputs from work I know well. They include agent paths from my navigation thesis and earth observation models published with IEEE and Remote Sensing of Environment. I also use open Southeast Asian language models built with SEACrowd and SEA-VL. English-first benchmarks rarely cover these inputs.

Go deeper

Review the complete professional record.

Notes on calibration, alignment training, and evaluation. Reading notes, experiment logs, and the occasional essay. I post the results that went against me too.

Knowing when you don't know is the core safety property

Why safe deployment depends on models knowing when to abstain.

Read the essay →
Featured in The Business Times

AI rules in SEA: the risks, the fines, what you need to know

Featured coverage of new AI rules, enforcement risks, and compliance duties across Southeast Asia.

Read in The Business Times →

Let's Collaborate

I study calibration under safety alignment and ways to keep safe models useful. If your work overlaps with the areas below, I would be glad to hear from you.

Research Collaboration Research & Applied Roles Speaking Mentorship
Work With Me →