当前位置：文江博客话题详情

如何从 SAS URL 访问方法中删除 HTML？

发布于 2024-07-23 08:56:16 字数 46 浏览 9 评论 0原文

使用SAS URL访问方法读取网页时，删除所有HTML标签的最便捷方法是什么？

原文

分享到QQ

分享到微博

如果你对这篇内容有疑问，欢迎到本站社区发帖提问参与讨论，获取更多帮助，或者扫码二维码加入 Web 技术交流群。

发布评论

需要登录才能够评论，你可以免费注册一个本站的账号。

奶气 2024-07-30 08:56:16

这应该做你想做的。删除 <> 之间的所有内容包括<> 并只留下内容（又名innerHTML）。

Data HTMLData;

filename INDEXIN URL "http://www.zug.com/";

input;

textline = _INFILE_;

/*-- Clear out the HTML text --*/
re1 = prxparse("s/<(.|\n)*?>//");
call prxchange(re1, -1, textline);

run;

This should do what you want. Removes everything between the <> including the <> and leaves just the content (aka innerHTML).

Data HTMLData;

filename INDEXIN URL "http://www.zug.com/";

input;

textline = _INFILE_;

/*-- Clear out the HTML text --*/
re1 = prxparse("s/<(.|\n)*?>//");
call prxchange(re1, -1, textline);

run;

回复收藏 0 原文