Home
cd ../playbooks
Academic ResearchBeginner

Scientific Gene Database

Query NCBI Gene via E-utilities/Datasets API. Search by symbol/ID, retrieve gene info (RefSeqs, GO, locations, phenotypes), batch lookups, for gene annotation and functional analysis.

5 minutes
By K-Dense AISource
#scientific#claude-code#gene-database#database#protein#genomics

You have a list of 200 gene symbols and need RefSeq IDs, chromosomal locations, GO annotations, and phenotype associations for all of them. Manually looking up each gene on NCBI takes hours. Programmatic access via E-utilities lets you batch-query gene information and build annotation tables in minutes.

Who it's for: bioinformaticians annotating gene lists from expression or GWAS studies, genomics researchers retrieving comprehensive gene information for candidate genes, clinical geneticists looking up gene-phenotype associations, functional genomics teams building gene annotation databases, graduate students learning to work with NCBI programmatic tools

Example

"Annotate our list of 150 differentially expressed genes with NCBI Gene data" → Gene annotation table: gene symbols mapped to Entrez IDs and RefSeq accessions, chromosomal locations, GO term annotations (biological process, molecular function), associated phenotypes and diseases, and a consolidated spreadsheet-ready annotation matrix

CLAUDE.md Template

New here? 3-minute setup guide → | Already set up? Copy the template below.

# Gene Database

## Overview

NCBI Gene is a comprehensive database integrating gene information from diverse species. It provides nomenclature, reference sequences (RefSeqs), chromosomal maps, biological pathways, genetic variations, phenotypes, and cross-references to global genomic resources.

## When to Use This Skill

This skill should be used when working with gene data including searching by gene symbol or ID, retrieving gene sequences and metadata, analyzing gene functions and pathways, or performing batch gene lookups.

## Quick Start

NCBI provides two main APIs for gene data access:

1. **E-utilities** (Traditional): Full-featured API for all Entrez databases with flexible querying
2. **NCBI Datasets API** (Newer): Optimized for gene data retrieval with simplified workflows

Choose E-utilities for complex queries and cross-database searches. Choose Datasets API for straightforward gene data retrieval with metadata and sequences in a single request.

## Common Workflows

### Search Genes by Symbol or Name

To search for genes by symbol or name across organisms:

1. Use the `scripts/query_gene.py` script with E-utilities ESearch
2. Specify the gene symbol and organism (e.g., "BRCA1 in human")
3. The script returns matching Gene IDs

Example query patterns:
- Gene symbol: `insulin[gene name] AND human[organism]`
- Gene with disease: `dystrophin[gene name] AND muscular dystrophy[disease]`
- Chromosome location: `human[organism] AND 17q21[chromosome]`

### Retrieve Gene Information by ID

To fetch detailed information for known Gene IDs:

1. Use `scripts/fetch_gene_data.py` with the Datasets API for comprehensive data
2. Alternatively, use `scripts/query_gene.py` with E-utilities EFetch for specific formats
3. Specify desired output format (JSON, XML, or text)

The Datasets API returns:
- Gene nomenclature and aliases
- Reference sequences (RefSeqs) for transcripts and proteins
- Chromosomal location and mapping
- Gene Ontology (GO) annotations
- Associated publications

### Batch Gene Lookups

For multiple genes simultaneously:

1. Use `scripts/batch_gene_lookup.py` for efficient batch processing
2. Provide a list of gene symbols or IDs
3. Specify the organism for symbol-based queries
4. The script handles rate limiting automatically (10 requests/second with API key)

This workflow is useful for:
- Validating gene lists
- Retrieving metadata for gene panels
- Cross-referencing gene identifiers
- Building gene annotation tables

### Search by Biological Context

To find genes associated with specific biological functions or phenotypes:

1. Use E-utilities with Gene Ontology (GO) terms or phenotype keywords
2. Query by pathway names or disease associations
3. Filter by organism, chromosome, or other attributes

Example searches:
- By GO term: `GO:0006915[biological process]` (apoptosis)
- By phenotype: `diabetes[phenotype] AND mouse[organism]`
- By pathway: `insulin signaling pathway[pathway]`

### API Access Patterns

**Rate Limits:**
- Without API key: 3 requests/second for E-utilities, 5 requests/second for Datasets API
- With API key: 10 requests/second for both APIs

**Authentication:**
Register for a free NCBI API key at https://www.ncbi.nlm.nih.gov/account/ to increase rate limits.

**Error Handling:**
Both APIs return standard HTTP status codes. Common errors include:
- 400: Malformed query or invalid parameters
- 429: Rate limit exceeded
- 404: Gene ID not found

Retry failed requests with exponential backoff.

## Script Usage

### query_gene.py

Query NCBI Gene using E-utilities (ESearch, ESummary, EFetch).

```bash
python scripts/query_gene.py --search "BRCA1" --organism "human"
python scripts/query_gene.py --id 672 --format json
python scripts/query_gene.py --search "insulin[gene] AND diabetes[disease]"
```

### fetch_gene_data.py

Fetch comprehensive gene data using NCBI Datasets API.

```bash
python scripts/fetch_gene_data.py --gene-id 672
python scripts/fetch_gene_data.py --symbol BRCA1 --taxon human
python scripts/fetch_gene_data.py --symbol TP53 --taxon "Homo sapiens" --output json
```

### batch_gene_lookup.py

Process multiple gene queries efficiently.

```bash
python scripts/batch_gene_lookup.py --file gene_list.txt --organism human
python scripts/batch_gene_lookup.py --ids 672,7157,5594 --output results.json
```

## API References

For detailed API documentation including endpoints, parameters, response formats, and examples, refer to:

- `references/api_reference.md` - Comprehensive API documentation for E-utilities and Datasets API
- `references/common_workflows.md` - Additional examples and use case patterns

Search these references when needing specific API endpoint details, parameter options, or response structure information.

## Data Formats

NCBI Gene data can be retrieved in multiple formats:

- **JSON**: Structured data ideal for programmatic processing
- **XML**: Detailed hierarchical format with full metadata
- **GenBank**: Sequence data with annotations
- **FASTA**: Sequence data only
- **Text**: Human-readable summaries

Choose JSON for modern applications, XML for legacy systems requiring detailed metadata, and FASTA for sequence analysis workflows.

## Best Practices

1. **Always specify organism** when searching by gene symbol to avoid ambiguity
2. **Use Gene IDs** for precise lookups when available
3. **Batch requests** when working with multiple genes to minimize API calls
4. **Cache results** locally to reduce redundant queries
5. **Include API key** in scripts for higher rate limits
6. **Handle errors gracefully** with retry logic for transient failures
7. **Validate gene symbols** before batch processing to catch typos

## Resources

This skill includes:

### scripts/
- `query_gene.py` - Query genes using E-utilities (ESearch, ESummary, EFetch)
- `fetch_gene_data.py` - Fetch gene data using NCBI Datasets API
- `batch_gene_lookup.py` - Handle multiple gene queries efficiently

### references/
- `api_reference.md` - Detailed API documentation for both E-utilities and Datasets API
- `common_workflows.md` - Examples of common gene queries and use cases
README.md

What This Does

NCBI Gene is a comprehensive database integrating gene information from diverse species. It provides nomenclature, reference sequences (RefSeqs), chromosomal maps, biological pathways, genetic variations, phenotypes, and cross-references to global genomic resources.


Quick Start

Step 1: Create a Project Folder

mkdir -p ~/Projects/gene-database

Step 2: Download the Template

Click Download above, then:

mv ~/Downloads/CLAUDE.md ~/Projects/gene-database/

Step 3: Start Claude Code

cd ~/Projects/gene-database
claude

Common Workflows

Search Genes by Symbol or Name

To search for genes by symbol or name across organisms:

  1. Use the scripts/query_gene.py script with E-utilities ESearch
  2. Specify the gene symbol and organism (e.g., "BRCA1 in human")
  3. The script returns matching Gene IDs

Example query patterns:

  • Gene symbol: insulin[gene name] AND human[organism]
  • Gene with disease: dystrophin[gene name] AND muscular dystrophy[disease]
  • Chromosome location: human[organism] AND 17q21[chromosome]

Retrieve Gene Information by ID

To fetch detailed information for known Gene IDs:

  1. Use scripts/fetch_gene_data.py with the Datasets API for comprehensive data
  2. Alternatively, use scripts/query_gene.py with E-utilities EFetch for specific formats
  3. Specify desired output format (JSON, XML, or text)

The Datasets API returns:

  • Gene nomenclature and aliases
  • Reference sequences (RefSeqs) for transcripts and proteins
  • Chromosomal location and mapping
  • Gene Ontology (GO) annotations
  • Associated publications

Batch Gene Lookups

For multiple genes simultaneously:

  1. Use scripts/batch_gene_lookup.py for efficient batch processing
  2. Provide a list of gene symbols or IDs
  3. Specify the organism for symbol-based queries
  4. The script handles rate limiting automatically (10 requests/second with API key)

This workflow is useful for:

  • Validating gene lists
  • Retrieving metadata for gene panels
  • Cross-referencing gene identifiers
  • Building gene annotation tables

Search by Biological Context

To find genes associated with specific biological functions or phenotypes:

  1. Use E-utilities with Gene Ontology (GO) terms or phenotype keywords
  2. Query by pathway names or disease associations
  3. Filter by organism, chromosome, or other attributes

Example searches:

  • By GO term: GO:0006915[biological process] (apoptosis)
  • By phenotype: diabetes[phenotype] AND mouse[organism]
  • By pathway: insulin signaling pathway[pathway]

API Access Patterns

Rate Limits:

  • Without API key: 3 requests/second for E-utilities, 5 requests/second for Datasets API
  • With API key: 10 requests/second for both APIs

Authentication: Register for a free NCBI API key at https://www.ncbi.nlm.nih.gov/account/ to increase rate limits.

Error Handling: Both APIs return standard HTTP status codes. Common errors include:

  • 400: Malformed query or invalid parameters
  • 429: Rate limit exceeded
  • 404: Gene ID not found

Retry failed requests with exponential backoff.

Script Usage

query_gene.py

Query NCBI Gene using E-utilities (ESearch, ESummary, EFetch).

python scripts/query_gene.py --search "BRCA1" --organism "human"
python scripts/query_gene.py --id 672 --format json
python scripts/query_gene.py --search "insulin[gene] AND diabetes[disease]"

fetch_gene_data.py

Fetch comprehensive gene data using NCBI Datasets API.

python scripts/fetch_gene_data.py --gene-id 672
python scripts/fetch_gene_data.py --symbol BRCA1 --taxon human
python scripts/fetch_gene_data.py --symbol TP53 --taxon "Homo sapiens" --output json

batch_gene_lookup.py

Process multiple gene queries efficiently.

python scripts/batch_gene_lookup.py --file gene_list.txt --organism human
python scripts/batch_gene_lookup.py --ids 672,7157,5594 --output results.json

API References

For detailed API documentation including endpoints, parameters, response formats, and examples, refer to:

  • references/api_reference.md - Comprehensive API documentation for E-utilities and Datasets API
  • references/common_workflows.md - Additional examples and use case patterns

Search these references when needing specific API endpoint details, parameter options, or response structure information.

Data Formats

NCBI Gene data can be retrieved in multiple formats:

  • JSON: Structured data ideal for programmatic processing
  • XML: Detailed hierarchical format with full metadata
  • GenBank: Sequence data with annotations
  • FASTA: Sequence data only
  • Text: Human-readable summaries

Choose JSON for modern applications, XML for legacy systems requiring detailed metadata, and FASTA for sequence analysis workflows.

Best Practices

  1. Always specify organism when searching by gene symbol to avoid ambiguity
  2. Use Gene IDs for precise lookups when available
  3. Batch requests when working with multiple genes to minimize API calls
  4. Cache results locally to reduce redundant queries
  5. Include API key in scripts for higher rate limits
  6. Handle errors gracefully with retry logic for transient failures
  7. Validate gene symbols before batch processing to catch typos

Resources

This skill includes:

scripts/

  • query_gene.py - Query genes using E-utilities (ESearch, ESummary, EFetch)
  • fetch_gene_data.py - Fetch gene data using NCBI Datasets API
  • batch_gene_lookup.py - Handle multiple gene queries efficiently

references/

  • api_reference.md - Detailed API documentation for both E-utilities and Datasets API
  • common_workflows.md - Examples of common gene queries and use cases

$Related Playbooks

Academic Research

Scientific Geniml

This skill should be used when working with genomic interval data (BED files) for machine learning tasks. Use for training region embeddings (Region2Vec, BEDspace), single-cell ATAC-seq analysis (scEmbed), building consensus peaks (universes), or ...

10 minutes
Intermediate
Academic Research

Scientific Geo Database

Access NCBI GEO for gene expression/genomics data. Search/download microarray and RNA-seq datasets (GSE, GSM, GPL), retrieve SOFT/Matrix files, for transcriptomics and expression analysis.

5 minutes
Beginner
Academic Research

Scientific Gget

'Fast CLI/Python queries to 20+ bioinformatics databases. Use for quick lookups: gene info, BLAST searches, AlphaFold structures, enrichment analysis. Best for interactive exploration, simple queries. For batch processing or advanced BLAST use bio...

10 minutes
Intermediate
Academic Research

Scientific Ginkgo Cloud Lab

Submit and manage protocols on Ginkgo Bioworks Cloud Lab (cloud.ginkgo.bio), a web-based interface for autonomous lab execution on Reconfigurable Automation Carts (RACs). Use when the user wants to run cell-free protein expression (validation or o...

10 minutes
Intermediate
Academic Research

Scientific Glycoengineering

Analyze and engineer protein glycosylation. Scan sequences for N-glycosylation sequons (N-X-S/T), predict O-glycosylation hotspots, and access curated glycoengineering tools (NetOGlyc, GlycoShield, GlycoWorkbench). For glycoprotein engineering, th...

15 minutes
Advanced
Academic Research

Scientific Gnomad Database

Query gnomAD (Genome Aggregation Database) for population allele frequencies, variant constraint scores (pLI, LOEUF), and loss-of-function intolerance. Essential for variant pathogenicity interpretation, rare disease genetics, and identifying loss...

5 minutes
Beginner
Academic Research

Scientific Gtars

High-performance toolkit for genomic interval analysis in Rust with Python bindings. Use when working with genomic regions, BED files, coverage tracks, overlap detection, tokenization for ML models, or fragment analysis in computational genomics a...

10 minutes
Intermediate
Academic Research

Scientific Gtex Database

Query GTEx (Genotype-Tissue Expression) portal for tissue-specific gene expression, eQTLs (expression quantitative trait loci), and sQTLs. Essential for linking GWAS variants to gene regulation, understanding tissue-specific expression, and interp...

5 minutes
Beginner
Academic Research

Scientific Gwas Database

Query NHGRI-EBI GWAS Catalog for SNP-trait associations. Search variants by rs ID, disease/trait, gene, retrieve p-values and summary statistics, for genetic epidemiology and polygenic risk scores.

5 minutes
Beginner
Academic Research

Scientific Histolab

Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatia...

10 minutes
Intermediate
Academic Research

Scientific Hmdb Database

Access Human Metabolome Database (220K+ metabolites). Search by name/ID/structure, retrieve chemical properties, biomarker data, NMR/MS spectra, pathways, for metabolomics and identification.

5 minutes
Beginner
Academic Research

Scientific Imaging Data Commons

Query and download public cancer imaging data from NCI Imaging Data Commons using idc-index. Use for accessing large-scale radiology (CT, MR, PET) and pathology datasets for AI training or research. No authentication required. Query by metadata, v...

15 minutes
Advanced

Browse all Academic Research playbooks →