Instructions to use panzs19/ContextPilot-E4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use panzs19/ContextPilot-E4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="panzs19/ContextPilot-E4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("panzs19/ContextPilot-E4B") model = AutoModelForMultimodalLM.from_pretrained("panzs19/ContextPilot-E4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use panzs19/ContextPilot-E4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "panzs19/ContextPilot-E4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "panzs19/ContextPilot-E4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/panzs19/ContextPilot-E4B
- SGLang
How to use panzs19/ContextPilot-E4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "panzs19/ContextPilot-E4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "panzs19/ContextPilot-E4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "panzs19/ContextPilot-E4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "panzs19/ContextPilot-E4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use panzs19/ContextPilot-E4B with Docker Model Runner:
docker model run hf.co/panzs19/ContextPilot-E4B
ContextPilot-E4B
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
ContextPilot-E4B is the Gemma4-E4B checkpoint of ContextPilot, a proactive context-management framework for long-horizon language-model agents. It teaches agents to plan, maintain long-term memory, and offload less useful context while they continue reasoning and using tools. For more details, see our paper and code repository.
Overview
ContextPilot combines three main components:
- an extended context-management toolset with planning, structured memory, retrieval, and soft context offloading;
- context-aware partial rollout that focuses exploration on sensitive context-editing decisions; and
- fine-grained credit assignment that trains intermediate snapshots using the outcomes of their downstream branches.
The resulting agents are evaluated on long-context question answering and deep-search tasks; see the evaluation instructions for details.
Loading
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "panzs19/ContextPilot-E4B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
Note that loading the checkpoint alone does not execute context-management tools; the tool definitions, agent runtime, and evaluation pipeline are provided in the ContextPilot repository. See the inference guide for the full setup.
Intended Use
This checkpoint is intended for research on proactive context management, long-horizon agents, long-context QA, and deep search.
- Downloads last month
- 14
