datasets_journalists_search
Search the journalists dataset. Searches the journalists index (dataset id enum value journalists) — public journalist and reporter contact records crawled from news outlets' own staff/author pages, for PR outreach. Each record carries the outlet, title, best-effort beat topics, and any public co...
This record as markdown: /tools/crawlora-mcp/datasets-journalists-search.md
What datasets_journalists_search does on Crawlora
AI agents call datasets_journalists_search to retrieve information from Crawlora without modifying anything. It is typically the context-gathering step in research, monitoring, and reporting workflows, before the agent takes action elsewhere.
| Parameter | Type | Required | Description |
|---|---|---|---|
q | string | — | Full-text match on the journalist's name, title, and bio, max 256 characters |
page | integer | — | Page number, defaults to 1 |
sort | string | — | Sort enum: relevance, name_asc, outlet_asc, crawled_desc |
topic | string | — | Exact topic filter, e.g. security, stablecoins. Use the values returned by facets?facet=topic |
outlet | string | — | Exact outlet id filter, e.g. techcrunch, coindesk. Use the ids returned by facets?facet=outlet |
vertical | string | — | Exact beat-vertical filter. Enum: tech, crypto, marketing, consumer_tech, consumer_policy, cybersecurity, health, gaming, climate, business, entertainment, spor |
page_size | integer | — | Page size, defaults to 20 and maxes at 100; page * page_size must be <= 10000 |
contact_type | string | — | Contact-availability filter. Enum: email, social, none |
Parameters from the server's own tool schema.
Why datasets_journalists_search is rated Low
Even though datasets_journalists_search only reads data, uncontrolled read access leaks sensitive information and racks up API costs: an agent caught in a retry loop can make thousands of calls a minute without anyone noticing.
Attacks that exploit this kind of access
The rule that runs datasets_journalists_search safely
PolicyLayer is an MCP gateway: it sits between your AI agents and Crawlora, and checks every tool call against a rule you set before the call runs. Nothing changes on the server itself. For datasets_journalists_search, this is the rule to start with:
datasets_journalists_search is read-only, so it stays allowed. Everything else on the server is denied unless you say otherwise.
The button opens the PolicyLayer dashboard: create your workspace, connect Crawlora, apply this rule, and every datasets_journalists_search call is checked against it from then on.
Questions about datasets_journalists_search
Search the journalists dataset. Searches the journalists index (dataset id enum value journalists) — public journalist and reporter contact records crawled from news outlets' own staff/author pages, for PR outreach. Each record carries the outlet, title, best-effort beat topics, and any public contact info (a work email or a social handle) found on that outlet's own page. There is no cross-outlet upstream search; this dataset is built by crawling a curated roster of outlets ourselves. vertical enum: tech, crypto, marketing, consumer_tech, consumer_policy, cybersecurity, health, gaming, climate, business, entertainment, sports, legal, science, politics, real_estate, automotive, travel, food, education, design, film_tv, fashion, music, personal_finance, tech_independent, culture_independent, local_news, construction, banking, retail, aerospace_defense, energy, agriculture, local_business. contact_type enum: email, social, none. sort enum: relevance, name_asc, outlet_asc, crawled_desc. It is categorised as a Read tool in the Crawlora MCP Server, which means it retrieves data without modifying state.
datasets_journalists_search accepts 8 parameters: q, page, sort, topic, outlet, vertical, page_size, contact_type. The full parameter table on this page comes from the server's own tool schema.
Register the Crawlora MCP server in PolicyLayer and add a rule for datasets_journalists_search: allow, deny, rate-limit, or require approval. Point your MCP client at the PolicyLayer proxy URL and the rule is enforced on every call, before it reaches Crawlora. Nothing to install.
datasets_journalists_search is a Read tool with low risk. Read-only tools are generally safe to allow by default.
Yes. Add a rate_limit block to the datasets_journalists_search rule in your PolicyLayer policy. For example, setting max: 10 and window: 60 limits the tool to 10 calls per minute. Rate limits are tracked per agent session and reset automatically.
Set action: deny in the PolicyLayer policy for datasets_journalists_search. The AI agent will receive a policy violation error and cannot call the tool. You can also include a reason field to explain why the tool is blocked.
datasets_journalists_search is provided by the Crawlora MCP server (crawlora-mcp). PolicyLayer sits as a proxy in front of this server to enforce policies before tool calls reach the server.
More on Crawlora, and thousands of servers like it.
This server
Across the catalogue