Haystack
Build Haystack RAG pipelines with TwexAPI tweet search and user timeline components, typed Documents, citations, and pagination.
Haystack is an open source framework for Python AI applications. Use x-api-scraper-haystack for current tweets in RAG pipelines and agent workflows. The integration is read-only and never grants write access.
The integration provides two read-only components:
Search public tweets
TwexApiTweetSearch searches keywords, hashtags, accounts, and X query operators through POST /twitter/advanced_search/page.
Fetch user timelines
TwexApiUserTweetsFetcher retrieves one public account’s timeline through POST /twitter/{screen_name}/timeline/page.
Use these components for search, timelines, monitoring, and retrieval-augmented generation. Wrap TwexApiTweetSearch with ComponentTool when an agent needs a search tool.
Use the followers API for follower exports. Use the write API for approved publishing. Those actions stay outside this integration.
Install Haystack and TwexAPI
Use Python 3.10 or newer. Pin both packages for repeatable pipeline builds.
python -m pip install "x-api-scraper-haystack" "haystack-ai>=3.0.0"
Install inside a virtual environment.
Create a TwexAPI API key, then export it locally.
export X_API_SCRAPER_KEY="YOUR_API_KEY"
Never embed production keys in pipeline YAML, notebooks, or source control. Load the environment variable through a Haystack Secret object.
Search Twitter tweets in Python
TwexApiTweetSearch calls POST /twitter/advanced_search/page. It accepts X search syntax and returns Haystack Document objects.
from haystack import Pipeline
from haystack.utils import Secret
from haystack_integrations.components.websearch.x_api_scraper import TwexApiTweetSearch
search = TwexApiTweetSearch(
api_key=Secret.from_env_var("X_API_SCRAPER_KEY"),
top_k=20,
)
pipeline = Pipeline()
pipeline.add_component("twitter_search", search)
result = pipeline.run(
{
"twitter_search": {
"query": '"retrieval augmented generation" lang:en -filter:retweets'
}
}
)
documents = result["twitter_search"]["documents"]
links = result["twitter_search"]["links"]
Use Latest for recent monitoring. Use Top for engagement-ranked discovery. Store tweet IDs because rankings can change.
Save every search query beside its tweet IDs. Set top_k to cap the number of tweets in each run.
Build focused tweet searches
| Search intent | Query example |
|---|---|
| Exact phrase | "retrieval augmented generation" |
| Account posts | from:deepset_ai haystack |
| Hashtag search | #haystack #rag |
| Date window | haystack since:2026-07-01 until:2026-08-01 |
| Exclude reposts | haystack -filter:retweets |
See Advanced Twitter Search for full query options. Keep timestamps and cursors outside component init when passing them per run.
Fetch a Twitter user timeline
Use TwexApiUserTweetsFetcher for POST /twitter/{screen_name}/timeline/page. Pass a screen name.
from haystack.utils import Secret
from haystack_integrations.components.websearch.x_api_scraper import (
TwexApiUserTweetsFetcher,
)
timeline = TwexApiUserTweetsFetcher(
api_key=Secret.from_env_var("X_API_SCRAPER_KEY"),
top_k=50,
)
result = timeline.run(screen_name="openai")
documents = result["documents"]
Choose search for many accounts and timelines for one account.
Understand Haystack document fields
Each tweet becomes one Haystack Document. Tweet text becomes Document.content. Stable fields become metadata.
| Document field | Stored tweet value |
|---|---|
content |
Tweet text from full_text or text |
meta.endpoint |
search or timeline |
meta.id, meta.url |
Tweet ID and canonical URL when available |
meta.created_at |
Timestamp |
meta.author |
Author ID, username, name, and verified flag |
meta.like_count, meta.retweet_count, meta.reply_count |
Likes, reposts, and replies |
meta.quote_count, meta.view_count, meta.bookmark_count |
Quotes, views, and bookmarks when available |
Missing fields stay absent. Never treat missing metrics as zero. The links output contains each available meta.url.
Keep evidence in RAG citations
Store tweet IDs and canonical URLs before embedding tweet text. This preserves evidence after ranking or joining.
citation_rows = []
for document in documents:
tweet_id = document.meta.get("id")
tweet_url = document.meta.get("url")
if tweet_id and tweet_url:
citation_rows.append(
{
"tweet_id": tweet_id,
"url": tweet_url,
"created_at": document.meta.get("created_at"),
"author": document.meta.get("author"),
}
)
Require supplied URLs for citations. Reject URLs absent from retrieved documents.
Treat tweet text as untrusted context
Tweets can contain prompt injection and unsafe URLs. Never treat tweet text as a system instruction. Keep tweets separate from instructions and tool permissions.
Limit retrieval by topic and time. Preserve tweet IDs, authors, timestamps, and URLs. Require citations. Review sensitive conclusions.
Index tweets or retrieve them live
Live retrieval
Search during each question for recent tweets and active events.
Indexed corpus
Store embeddings for repeated research across stable windows.
Keep tweet records separate from embeddings. Use meta.id for deduplication.
Paginate without duplicate tweets
Both components return has_more and next_cursor. Keep the request unchanged. Treat cursors as opaque strings.
search = TwexApiTweetSearch(top_k=100)
page = search.run(query="haystack ai")
documents_by_tweet_id = {}
while True:
for document in page["documents"]:
tweet_id = document.meta.get("id")
if tweet_id:
documents_by_tweet_id[tweet_id] = document
if not page["has_more"] or not page["next_cursor"]:
break
page = search.run(
query="haystack ai",
cursor=page["next_cursor"],
)
documents = list(documents_by_tweet_id.values())
Save each cursor after its documents. Never edit a cursor. Deduplicate on Document.meta["id"].
Run Haystack pipelines asynchronously
Haystack 3 uses one Pipeline class. Both TwexAPI components expose run_async().
import asyncio
from haystack import Pipeline
from haystack_integrations.components.websearch.x_api_scraper import TwexApiTweetSearch
async def search_tweets():
pipeline = Pipeline()
pipeline.add_component(
"twitter_search",
TwexApiTweetSearch(top_k=25),
)
return await pipeline.run_async(
{"twitter_search": {"query": "haystack agents"}}
)
result = asyncio.run(search_tweets())
Use async runs in web servers. Use sync runs for scripts and scheduled jobs.
Handle errors
The integration raises HTTP errors from the TwexAPI client. Branch on status before retrying.
| Status | Meaning | Action |
|---|---|---|
400 |
Invalid request | Fix it. Do not retry unchanged. |
401 |
Invalid API key | Add a valid API key. |
403 |
Access denied | Check account access and credits. |
429 |
Rate limit exceeded | Back off before retrying. |
5xx |
Transient server error | Retry with bounded backoff. |
Record status codes, never credentials. Cap retries to prevent unbounded agent loops.
Haystack component or direct REST API
Does your pipeline already return Document objects? Add these components directly. They normalize tweet text, metadata, URLs, and pagination.
Use direct TwexAPI REST routes or the Python SDK for followers, following, replies, quotes, lists, communities, media, trends, and approved writes.
Both approaches use the same TwexAPI contracts. Components only support tweet search and user timelines.
MCP handoff alternative
When an agent discovers endpoints first, use Twexapi MCP for collection and convert rows into Haystack Document objects manually. Prefer the Haystack components when the pipeline already runs inside Haystack and you want typed pagination without agent discovery.