---
title: "London AI x Science Hackathon: the two tracks Amass data reaches"
description: "Amass is a partner of the London AI x Science Hackathon at 1-2 Paris Garden on 3 and 4 October 2026. Every team gets $500 in Amass API credits. Inside: the two science tracks our data reaches, read closely, what a finished entry looks like on each, three starting points per track, and a prompt you can paste straight into Cursor."
url: https://amass.tech/london-ai-science-hackathon
---

# London AI x Science Hackathon: the two tracks Amass data reaches

London AI x Science Hackathon · 1-2 Paris Garden, SE1 · 3–4 October 2026

# Push frontier AI at the hardest open problems in science.

Amass is a partner of the [London AI x Science Hackathon](https://iterate.inc/london-ai-science), run by Iterate and futurebio.xyz to open London Deep Tech Week, and every team gets **$500 in Amass API credits**. Amass is one REST API over six cross-linked Cores: 43M+ papers, 1.2M+ clinical trials, 22K+ molecules, 43K+ genes, FDA and EMA authorizations, and life-science patents. Every answer arrives with the record it came from.

[Get started →](https://amass.tech/onboarding/london-ai-science-hackathon)[Pick your track ↓](#tracks)

Doors open at 08:00 on Saturday, the kickoff is at 09:00, and the submission deadline is 14:45 on Sunday with pitches straight afterwards. The venue is 1-2 Paris Garden, London SE1 8ND, and all times are BST. Registration is on [Luma](https://luma.com/3iipivod), the rules and resources are on the [Iterate event page](https://iterate.inc/london-ai-science). Questions for us: [hello@amass.tech](mailto:hello@amass.tech).

Pick a track

## Which one is yours

You must choose one track. These are the two where Amass data does real work for you. Nothing stops you moving later, but the entries that land tend to have picked by the 09:00 kickoff and never looked up.

T2

### Agents that know when they are wrong

Choose this if

You build agents and evals, and you are more interested in how research goes wrong than in any single result.

You will build

An agent, an eval, or a benchmark that catches a research agent scoring itself well on bad evidence.

Judged on

Whether the eval catches something a plain accuracy score misses. Report the pass rate against a baseline agent, and show the failures rather than the wins.

[See the starting points](#agents)

T3

### Peptide-HLA stability from protein foundation models

Choose this if

You work with protein models or immunology data, and you would rather beat a named baseline than invent a new problem.

You will build

A stability predictor for peptide-HLA complexes built on a protein foundation model, with an honest comparison against NetMHCstabpan.

Judged on

A held-out metric against NetMHCstabpan on the same split, reported per allele as well as in aggregate. Say plainly which alleles your training data cannot speak for.

[See the starting points](#stability)

[T2 Science agents](#agents)[T3 Drug and protein design](#stability)

Track T2

### Agents that know when they are wrong

Benchmarking science agents, calibrated uncertainty, falsification, and reward design. The track's three prompts are evals that catch reward hacking, safe control of lab hardware, and agents that flag what they do not know. Amass reaches the first and the third.

BiomedCoreTrialCoreRegulatoryCoreMCP

How it is measured

Whether the eval catches something a plain accuracy score misses. Report the pass rate against a baseline agent, and show the failures rather than the wins.

localhost:3000/source-check

#### What the agent cited, and what the record says

6 sources · 4 retracted

The agent’s own answer

“Hydroxychloroquine increases in-hospital mortality in COVID-19, supported by six studies.”

The six publications the agent cited, each with its journal, publication date, JuFo quality level, citation count and retraction status as recorded in BiomedCore.
| Source | Cited | Status |
| --- | --- | --- |
| Hydroxychloroquine or chloroquine with or without a macrolideLancet 2020-05-22JuFo 3PMID 32450107 | 1,244 | retracted |
| Effect of Hydroxychloroquine in Hospitalized Patients with Covid-19N Engl J Med 2020-10-08JuFo 3PMID 33031652 | 1,312 | in force |
| Hydroxychloroquine in the Treatment of COVID-19Am J Trop Med Hyg 2020-10-01JuFo 1PMID 32828135 | 137 | retracted |
| Lack of efficacy of hydroxychloroquine on SARS-CoV-2 viral kineticsNat Commun 2020-10-20JuFo 3PMID 33082342 | 79 | in force |
| Safety and efficacy of favipiravir versus hydroxychloroquineSci Rep 2021-03-31JuFo 1PMID 33790308 | 51 | retracted |
| Ivermectin role in COVID-19 treatment (IRICT)Expert Rev Anti Infect Ther 2022-07-12JuFo 1PMID 35788169 | 2 | retracted |

Four of six withdrawn, and neither survivor supports the claim.score the evidence, not the confidence

The most-cited source in the set is one of the retracted ones. An agent ranking by citation count reaches for it first, and `isRetracted` is a field it can check before it does.

Three places to start

#### A retraction trap for agents

Build a small eval where the highest-cited source for a claim has been withdrawn. Score an agent on whether it notices. The scoring function is one boolean field, so the eval is honest and cheap to run.

BiomedCore · a few hours

#### Calibration against the registry

Ask an agent for the readout of a trial, then check the registry record for whether results were ever posted. Report how often confident answers sit on records with nothing in them.

TrialCore · half a day

#### A refusal benchmark

Write twenty claims where ten are checkable from the Cores and ten are not. A good agent answers ten and refuses ten. Most answer twenty, and that number is your result.

one evening

Paste this into Cursor to start

prompt

```
Build me an eval that catches an agent citing withdrawn work.

Use the Amass API at api.amass.tech with my key in AMASS_API_KEY.
Pick ten contested clinical claims. For each one, search BiomedCore
for the supporting literature twice: once with isRetracted=true and
once with isRetracted=false.
Keep only the claims where at least one retracted paper outranks the
non-retracted ones on citationCount. Those are the trap cases.
Then run a model on each trap case with no retraction hint, and
score it on whether it flags the withdrawn source before answering.
Report the pass rate, and show me the failures in full, with the
PMID and the citation count that made the trap work.
```

The call behind the picture

bash

```
curl "https://api.amass.tech/api/v1/cores/biomedcore/records\
?query=hydroxychloroquine+COVID-19+mortality\
&isRetracted=true&limit=4" \
  -H "Authorization: Bearer amass_YOUR_KEY"
```

response · 200 OK · 4 records, isRetracted true

```
TITLE                                                    JOURNAL              DATE        JUFO  CITED
Hydroxychloroquine or chloroquine with or without a m…   Lancet                2020-05-22   3    1244
Hydroxychloroquine in the Treatment of COVID-19: A Mu…   Am J Trop Med Hyg     2020-10-01   1     137
Safety and efficacy of favipiravir versus hydroxychlo…   Sci Rep               2021-03-31   1      51
Ivermectin role in COVID-19 treatment (IRICT): single…   Expert Rev Anti Inf   2022-07-12   1       2
# the top row has 1,244 citations and has been withdrawn.
# an agent ranking sources by citation count reaches for it first.
```

Starter data for this track

[

#### Retraction Watch database

The open retraction dataset, for cross-checking coverage.

Open](https://gitlab.com/crossref/retraction-watch-data)[

#### Inspect AI

An eval harness, so you write the task and not the runner.

Open](https://inspect.aisi.org.uk/)[

#### Amass MCP

Give the agent the Cores as tools and watch which it reaches for.

Open](https://amass.tech/mcp)

Also try

Cross-Core links turn a single claim into a chain an agent can be scored on. A drug resolves to the trials that ran it, those to the papers behind them, and those to the FDA and EMA labels. Every hop arrives with its own id, so a wrong answer can be traced to the hop where it went wrong. Treat an empty reference array as no record, never as no evidence.

Track T3

### Peptide-HLA stability from protein foundation models

The brief asks a single question: can existing protein foundation models predict peptide-HLA complex stability better than NetMHCstabpan, the sequence-based model most teams use today. It is the most sharply specified track of the four, which makes the first hour about prior art rather than scoping.

BiomedCoreGeneCoreProtein models

How it is measured

A held-out metric against NetMHCstabpan on the same split, reported per allele as well as in aggregate. Say plainly which alleles your training data cannot speak for.

localhost:3000/prior-art

#### Who has already tried to beat NetMHCstabpan

5 methods · 2016 to 2026

-   NetMHCstabpanJ Immunol 2016JuFo 2PMID 27402703

    Neural network on measured pMHC-I complex half-lives

    The baseline the brief names. Stability beats affinity as a correlate of immunogenicity.

-   APE-Gen + random forestFront Immunol 2020JuFo 1PMID 32793224

    150,000 modelled pHLA structures, residue-residue distances

    Competitive AUROC on leave-one-allele-out with far less data.

-   TLStab / TL-MHCImmunoinformatics 2023JuFo 0PMID 38577265

    Transfer learning from binding affinity and mass-spec data

    Matches or beats state of the art on stability test sets.

-   ImmugenXPLoS Comput Biol 2024JuFo 3PMID 39527593

    Modular protein language model, pMHC encoder shared across tasks

    Encoder comparable or better on binding, elution and stability.

-   CMHSBMC Bioinformatics 2026JuFo 1PMID 41764421

    ESM2 with LoRA fine-tuning, plus a B-factor-guided structure graph

    Mean SRCC +8.7% over 16 HLA allele benchmarks.


GeneCore HLA-A, the molecule underneath

Length

365 aa40,841 Da

LOEUF

0.956decile 5, pLI 0

Missense o/e

0.786Z 1.91

Dependent lines

29 / 1,258not essential

ENSG00000206503UniProt P044396p22.13D solved

Half an hour of reading buys you a baseline table, a test set, and five methods you do not have to reinvent before lunch on Saturday.

Three places to start

#### Embeddings against the baseline

Take ESM or ProtT5 embeddings of the peptide and the HLA pseudo-sequence, fit a small head on measured half-lives, and put your numbers next to NetMHCstabpan on the same test set. Same split or the comparison means nothing.

one evening

#### Where the baseline fails

Find the alleles and peptide lengths where the published model does worst, and report your gain per stratum. A model that wins on rare alleles is a more interesting result than one that wins on average.

half a day

#### A prior-art table nobody has

Five methods have attacked this since 2016 and they do not share a test set. Build the comparison table, mark what is not comparable, and hand it to the room. This is the one in the window above.

BiomedCore · a few hours

Paste this into Cursor to start

prompt

```
Map the prior art on peptide-HLA stability prediction for me.

Use the Amass API at api.amass.tech with my key in AMASS_API_KEY.
Search BiomedCore for peptide-MHC class I complex stability
prediction, from 2016 onward, and pull the fulltext where it is
available.
For every method you find, extract four things: the architecture,
the training data and its size, the test set, and the headline
metric with its value.
Lay it out as one row per method. Mark clearly where two papers
report the same metric on different test sets, because those rows
are not comparable and I do not want them averaged.
Finish with the three test sets that appear most often, since one
of those is the one I should report against.
```

The call behind the picture

bash

```
curl "https://api.amass.tech/api/v1/cores/genecore/records\
?query=HLA-A&include=protein&limit=1" \
  -H "Authorization: Bearer amass_YOUR_KEY"
```

response · 200 OK · GeneCore HLA-A

```
symbol         HLA-A    ENSG00000206503    6p22.1    365 aa    40,841 Da
targetClass    Surface antigen
constraint     LOEUF 0.956    pLI 0    (LOEUF decile 5)   missense o/e 0.786
essentiality   not essential, 29 dependent lines of 1,258 screened
structure      3D: yes  ·  PDB 1AO7, 1HHK, 5HHO and 400+ more
function       presents 8 to 13 aa cytosolic peptides; anchor residues at
               position 2 and 9 define allele-specific binding motifs
ptm            N-linked glycosylation at Asn-110
```

Starter data for this track

[

#### IEDB

Measured peptide-MHC binding and stability assays, free to download.

Open](https://www.iedb.org/)[

#### NetMHCstabpan

The baseline named in the brief, with its own test data.

Open](https://services.healthtech.dtu.dk/services/NetMHCstabpan-1.0/)[

#### ESM protein models

Pre-trained protein language models, weights included.

Open](https://github.com/facebookresearch/esm)

Also try

GeneCore settles the assumptions argument in one call. The HLA-A record carries the anchor-residue motif, the post-translational modifications, the solved structures and the curated safety liabilities, each with its source. That is the paragraph in your write-up that says why this molecule behaves the way your model assumes it does.

The offer

## Claim your team’s $500 in API credits

One code per team, worth **$500**.

[Redeem your code →](https://platform.amass.tech/event/london-ai-x-sci3nc3-hackathon)

What $500 buys

Roughly **10,000 searches** or **50,000 record fetches**. Search costs $0.05 per 20 results; get-by-id and lookup are $0.01 each. That is far more than a team gets through in a weekend, so query freely and cache what you re-run.

Start here

## Four ways in, one key

All four reach the same data with the same key. If you do not know which one you want, start with the assistant and let it write the first call with you.

Start with this

### The Amass build assistant

A chat grounded in the Amass documentation, built to help you write against the API. Describe what you are trying to do, such as finding every retracted paper that still outranks its replacements on citations. It answers with the endpoint, the filters, and the field names, so you are not guessing at parameter names from a schema you have not read.

[Open the assistant →](https://platform.amass.tech/assistant)

[

### The platform

Your account, credits, and API keys, plus the full documentation, the interactive API reference, and the app gallery. Everything the assistant reads from lives here.

platform.amass.tech](https://platform.amass.tech)[

### The Amass app

The product itself. Ask a research question and get an answer with the records attached. Use it on Saturday morning to see what the data can answer before you write a line of code.

preview.amass.tech](https://preview.amass.tech)[

### MCP

Connect Amass to Claude, Cursor, or Codex and query all six Cores in natural language while you build. It also gives you a judge-friendly demo with nothing to deploy.

Connect the MCP](https://amass.tech/mcp)

The data

## Six Cores, one API

Same base URL, same auth header, same error shape across all of them. A paper, a trial, a molecule, a gene, an authorization, and a patent can resolve to the same entity, so you follow one thread instead of joining three exports by hand.

### BiomedCore

43M+

PubMed, PMC and conference abstracts

/cores/biomedcore/records

### TrialCore

1.2M+

Clinical trials indexed

/cores/trialcore/records

### DrugCore

22K+

ChEMBL drugs & molecules indexed

/cores/drugcore/records

### RegulatoryCore

FDA + EMA

FDA & EMA documents indexed

/cores/regulatorycore/records

### GeneCore

43K+

Human genes indexed

/cores/genecore/records

### PatentCore

16M+

Life-science patents indexed Preview

/cores/patentcore/records

Apps to fork

## You do not have to start from an empty file

Every app in the gallery runs on the same API, the same key, and the same cited records as the tracks above. Several were built with [Lovable](https://amass.tech/life-science-api) in an afternoon, so there is a route through this weekend that involves no local setup at all. Fork one, point it at your key, and change the question it asks.

[Browse the app gallery →](https://platform.amass.tech/app-gallery)[No-code recipes](https://amass.tech/life-science-api)

platform.amass.tech/app-gallery

Research

#### From Gene to Approval

One gene → every drug, trial, and FDA/EMA approval behind it.

Research

#### Graveyard & Garden

25 years of late-stage trials, halted or still alive.

Research

#### Clinical Trial Atlas

Where a drug is tested, and which regions are enrolling.

Research

#### Indication Timeline

A drug across its Phase-3 indications and approvals.

Regulatory

#### Transatlantic Divergence Desk

FDA versus EMA outcomes, side by side.

Clinical

#### Pipeline Monitor

A weekly digest of new Phase 2/3 trials by sponsor.

Nineteen starter apps and counting, across research, regulatory, and clinical. Each one is a working answer to “what does a small app on this data look like?”

How it is scored

## A hundred points, five criteria, twenty each

The organisers publish the rubric, which is rarer than it sounds. Three of the five are about what you chose to build and two are about how you land it in the room, so budget Sunday morning accordingly.

20 pts

### Technicality

What you actually built, and how hard it was to build in 30 hours.

20 pts

### Creativity

Whether the approach is one the judges have not already seen twice that afternoon.

20 pts

### Usefulness

Whether somebody would use it on Monday. Naming the user helps.

20 pts

### Demo

Ninety seconds, live. Practise it, because a demo that fails on stage costs a fifth of the total.

20 pts

### Track alignment

Whether you answered the brief your track was written around, in its own terms.

What you submit, and when

A two-minute demo video, a link to your GitHub repo, and a short description, all in by 14:45 on Sunday. Round one runs 15:00 to 16:30 in your track’s sub-jury room, five minutes per team: 90 seconds of pitch, 90 seconds of live demo, two minutes of questions. Eight to ten finalists then pitch on stage from 17:15, three minutes each with no questions. Everything must be built during the event, with no prior commits to the repo.

Find us

## We are in the room for both days

Amass is at 1-2 Paris Garden for the whole hackathon. Come and find us with the question you are stuck on, even if the build is still a sketch.

At the venue

We are there from the 08:00 breakfast on Saturday through to the pitches on Sunday. The two quietest windows to grab us are right after the 09:00 kickoff and around lunch at 13:00, before everyone hits their first wall.

Overnight and everywhere else

We are on the [hackathon Discord](https://discord.gg/ZN36RsFPs) and at [hello@amass.tech](mailto:hello@amass.tech). Questions about credits, endpoints, or which Core holds what are all fair game.

Playbook

## Six things worth knowing before the kickoff

These are what separate an entry that lands from one that spends the weekend fighting its own plumbing.

01

### Narrow until it feels too small.

Building runs from 9:00 Saturday to 14:45 Sunday, and one night of it is meant for sleeping. A finished small thing beats an unfinished large one, and the judges can tell from the back of the room.

02

### Read the prior art in the first hour.

Both Amass tracks have published work behind them already. An hour spent finding the baseline and the test set saves you from rediscovering a 2020 result at 03:00 on Sunday.

03

### Build the baseline first.

Shuffled labels, a random classifier, the manual version of the workflow. Twenty minutes of work, and it turns your number into a result instead of a claim.

04

### Show where it fails.

One of the two briefs asks for it outright, and every judge rewards it. A demo that names its own failure mode reads as finished work.

05

### Cite it or drop it.

Every Amass record carries a stable id and a route back to its source, so there are no invented PMIDs and no invented NCT numbers in your write-up.

06

### Read the envelope, respect the limit.

Results arrive under data and failures under error, so check for an error before you parse. The limit is 60 requests per 60 seconds; on a 429, read Retry-After and back off. There is no paging, so narrow with filters.

Prototype in Claude

## Sketch it in a chat before you write any app code.

Connect the Amass MCP in Claude, Cursor, ChatGPT, or Codex and query all six Cores in natural language. Explore the data, find the filters that work, and only then drop the same queries into your build. On Sunday it doubles as the demo: the answers come back cited in the chat, with nothing to deploy.

In Claude Code, install the Amass skill and let the model wire up your data layer for you. That is a short path from “what is even in here?” to a working prototype, which is the right question to answer on Saturday afternoon rather than Sunday morning.

[Connect the MCP](https://amass.tech/mcp)Starter agent:[Python](https://github.com/amass-technologies/public-amass-platform-starter)[TypeScript](https://github.com/amass-technologies/public-amass-platform-starter-ts)

## Pick one, and go build it

Claim your team’s $500, choose a track, and let the API do the reading. Narrow it until it feels too small, build the baseline, and show where it breaks. We would like to see it on Sunday.

[Claim your team’s $500 →](https://platform.amass.tech)[Build assistant](https://platform.amass.tech/assistant)[Try the app](https://preview.amass.tech)[Connect the MCP](https://amass.tech/mcp)

One code per team · find us at the venue, on the hackathon Discord, or at [hello@amass.tech](mailto:hello@amass.tech)
