We build and train our own text-rewriting model: an open-weight LLM, fine-tuned and then improved with reinforcement learning. We are hiring a research engineer to run experiments on it.
This is a research support role. You are not setting the direction. You are running real experiments, building the training data they consume, and producing the measurements everything else gets decided on.
One thing that makes this job unusual, so you know up front: a lot of our evaluation and data labelling runs through paid third-party APIs. They are rate limited, occasionally flaky, and billed per run. Designing around that constraint is a real part of the work.
WHAT YOU WOULD BE DOING
1. Training data. Build and curate the corpora the model learns from - cleaning, format conversion, filtering, quality scoring, and generating matched pairs. This includes labelling passes that push large batches through external APIs. This is not grunt work; what goes into the dataset moves the model as much as the training recipe does.
2. Fine-tuning experiments. Parameter-efficient fine-tuning on open-weight bases, searching over dataset composition, base model, model size and training settings. Individual runs take a few hours. You would own running that grid properly, tracking it, and reporting what actually moved.
3. Reinforcement learning post-training. Policy-gradient RL on top of the fine-tuned model against a composite reward. You would run those experiments, and build, train and retrain the models the reward depends on as our targets change and more labelled data arrives.
4. Automated quality scoring. We score outputs on several dimensions using model-based graders, and calibrate those graders against expert human review. You would run that alignment loop - choosing grader models and prompts, and measuring how closely they track the human scores.
5. Evaluation and error analysis. Run the evaluation suites, keep results comparable across versions as the eval sets themselves change, and slice the results to find where the model actually fails - by document length, by document type, by failure mode. We also run automated checks over model output and rerun when they fire; measuring how often those checks fire correctly rather than wastefully is ongoing work. Adversarial prompt sets and statistical feature tests welcome. Nothing that changes model output ships without an evaluation behind it, and that gate would be yours.
REQUIREMENTS
Two hard requirements:
You have fine-tuned open-weight LLMs yourself, start to finish, with LoRA. Not through a managed fine-tuning API.
You can build an evaluation harness, keep it reproducible, and report results honestly, including when they are bad.
Then some of the following, and honesty about which:
Python, PyTorch and the Hugging Face stack (transformers, datasets, peft, trl). Unsloth or Axolotl a plus.
Working knowledge of RL post-training - GRPO, DPO or PPO - and of what goes wrong once a model starts optimizing against a metric.
Dataset engineering - cleaning, filtering, quality scoring, format conversion, building matched training pairs.
Training on rented GPUs, managing checkpoints and adapters on shared storage, and loading a model into an inference server so it can be evaluated. This is a small self-hosted setup, one training node, not a cluster.
NICE TO HAVE
Training small auxiliary models that feed into a training objective.
Containerized training and serving environments, and keeping an experiment reproducible across machines.
Multilingual work.
Experience working async in a small senior team.
HOW WE WORK
Async, written, fully remote, across several time zones. You run your own experiments and come back with results rather than being managed hour by hour. Evaluations cost real money per run, so we care about experiments being designed well rather than run often. If an evaluation says the new checkpoint is worse, we want to hear that the same day.
20 hours a week, ongoing, with room to grow. We do not care what hours you keep, but we need someone reachable across the week rather than in one block, and occasionally on a weekend when something genuinely needs it. Tell us your hourly rate.
In your first month you would take over running and reporting one experiment grid and one evaluation suite.
TO APPLY, ANSWER THESE THREE
1. Describe one open-weight model you fine-tuned end to end. Base model, method, dataset size, what you were optimizing for, and how you measured whether it worked. Give us one number from that work and the baseline you compared it against. A link to a repo, notebook or run log beats a description.
2. Tell us about a time a model you were training learned to exploit the metric you were scoring it on. What did the reward curve look like, what did the outputs look like, and what did you change? If it has never happened to you, say so plainly and tell us how you would catch it.
3. You rerun your evaluation on a new checkpoint and the score drops from 78% to 74% on a 200-example set. List, in order, what you check before telling anyone the model got worse - and tell us whether that difference is even worth reporting.
Please skip the generic proposal. Answer the three questions and we will reply.
You will leave ComplianceJobs.com.
We never charge candidates.
Job alerts
Get new jobs in your inbox