使用PHP解析XML

问题描述:

我一直在用PHP解析XML时遇到了一个问题,而且没有真正找到“正确的方式”或者至少是解析XML文件的标准方式。使用PHP解析XML

首先我试图分析此:

<item> 
    <title>2884400</title> 
    <description><![CDATA[ ><img width="126" alt="" src="http://userserve-ak.last.fm/serve/126/27319921.jpg" /> ]]></description> 
    <link>http://www.last.fm/music/+noredirect/Beatles/+images/27319921</link> 
    <author>anne710</author> 
    <pubDate>Tue, 21 Apr 2009 16:12:31 +0000</pubDate> 
    <guid>http://www.last.fm/music/+noredirect/Beatles/+images/27319921</guid> 
    <media:content url="http://userserve-ak.last.fm/serve/_/27319921/Beatles+2884400.jpg" fileSize="13065" type="image/jpeg" expression="full" width="126" height="126" /> 
    <media:thumbnail url="http://userserve-ak.last.fm/serve/126/27319921.jpg" type="image/jpeg" width="126" height="126" /> 
    </item> 

我使用这个代码:

$doc = new DOMDocument(); 
$doc->load('http://ws.audioscrobbler.com/2.0/artist/beatles/images.rss'); 
$arrFeeds = array(); 
foreach ($doc->getElementsByTagName('item') as $node) { 
    $itemRSS = array ( 
     'title' => $node->getElementsByTagName('title')->item(0)->nodeValue, 
     'desc' => $node->getElementsByTagName('description')->item(0)->nodeValue, 
     'link' => $node->getElementsByTagName('link')->item(0)->nodeValue, 
     'date' => $node->getElementsByTagName('pubDate')->item(0)->nodeValue 
     ); 
    array_push($arrFeeds, $itemRSS); 
} 

现在我想要得到的“媒体:内容”和“媒体:缩略图“网址属性,我会怎么做?现在我想我应该使用DOMElement :: getAttribute,但是我没有设法使它工作:/任何人都可以对此有所了解,并且让我知道这是否是解析XML的好方法?

问候, 沙迪

+0

这整个问题/线程是非常糟糕的。问题是缺乏对名字空间的理解。我建议任何人阅读这篇文章了解XML名称空间。人们在下面提到了这一点。问题在于media:内容意味着属于'media'命名空间的'content'标记,而不是默认的命名空间(这就是你要查询的内容)。 – Jotham 2010-02-11 02:04:18

这是如何我一直在使用的XMLReader最终做到了:

<?php 

define ('XMLFILE', 'http://ws.audioscrobbler.com/2.0/artist/vasco%20rossi/images.rss'); 
echo "<pre>"; 

$items = array(); 
$i = 0; 

$xmlReader = new XMLReader(); 
$xmlReader->open(XMLFILE, null, LIBXML_NOBLANKS); 

$isParserActive = false; 
$simpleNodeTypes = array ("title", "description", "media:title", "link", "author", "pubDate", "guid"); 

while ($xmlReader->read()) 
{ 
    $nodeType = $xmlReader->nodeType; 

    // Only deal with Beginning/Ending Tags 
    if ($nodeType != XMLReader::ELEMENT && $nodeType != XMLReader::END_ELEMENT) { continue; } 
    else if ($xmlReader->name == "item") { 
     if (($nodeType == XMLReader::END_ELEMENT) && $isParserActive) { $i++; } 
     $isParserActive = ($nodeType != XMLReader::END_ELEMENT); 
    } 

    if (!$isParserActive || $nodeType == XMLReader::END_ELEMENT) { continue; } 

    $name = $xmlReader->name; 

    if (in_array ($name, $simpleNodeTypes)) { 
     // Skip to the text node 
     $xmlReader->read(); 
     $items[$i][$name] = $xmlReader->value; 
    } else if ($name == "media:thumbnail") { 
     $items[$i]['media:thumbnail'] = array (
       "url" => $xmlReader->getAttribute("url"), 
       "width" => $xmlReader->getAttribute("width"), 
       "height" => $xmlReader->getAttribute("height"), 
       "type" => $xmlReader->getAttribute("type") 
     ); 
    } else if ($name == "media:content") { 
     $items[$i]['media:content'] = array (
       "url" => $xmlReader->getAttribute("url"), 
       "width" => $xmlReader->getAttribute("width"), 
       "height" => $xmlReader->getAttribute("height"), 
       "filesize" => $xmlReader->getAttribute("fileSize"), 
       "expression" => $xmlReader->getAttribute("expression") 
     ); 
    } 
} 

print_r($items); 
echo "</pre>"; 

?> 

你会想是这样的:

'content' => $node->getElementsByTagNameNS('http://search.yahoo.com/mrss/', 'content')->item(0)->getAttribute('url'); 
'thumbnail' => $node->getElementsByTagNameNS('http://search.yahoo.com/mrss/', 'thumbnail')->item(0)->getAttribute('url'); 

我相信会的工作,它已经有一段时间,因为我做了这样的事。

+0

那么如何实现呢? ! – 2009-07-13 20:57:06

+0

这是不是工作? – 2009-07-13 21:13:55

+0

[Mon Jul 13 23:13:04 2009] [error] [client xxx.xxx.xxx.xxx] PHP致命错误:在/ v2中的非对象上调用成员函数getAttribute()。73行上的php – 2009-07-13 21:21:34

<?php 

#Convert the String Into XML 
$xml = new SimpleXMLElement($_POST['name']); 

#Itterate through the XML for the data 

$values = "VALUES('' , "; 
foreach($xml->item as $item) 
{ 
//you now have access to that aitem 
} 

?> 

尝试使用SimpleXML:http://us2.php.net/simplexml

可以使用SimpleXML通过其他海报的建议,但你需要使用儿童()和属性()函数,所以你可以deal with the different namespaces

例(未经测试):

$feed = file_get_contents('http://ws.audioscrobbler.com/2.0/artist/beatles/images.rss'); 
$xml = new SimpleXMLElement($feed); 
foreach ($xml->channel->item as $item) { 
    foreach ($item->children('http://search.yahoo.com/mrss' as $media_element) { 
     var_dump($media_element); 
    } 
} 

或者,您可以使用XPath(再次,未经测试):

$feed = file_get_contents('http://ws.audioscrobbler.com/2.0/artist/beatles/images.rss'); 
$xml = new SimpleXMLElement($feed); 
$xml->registerXPathNamespace('media', 'http://ws.audioscrobbler.com/2.0/artist/beatles/images.rss'); 
$images = $xml->xpath('/rss/channel/item/media:[email protected]'); 
var_dump($images); 

试试这个。它会正常工作。

$doc = new DOMDocument(); 
$doc->load('http://ws.audioscrobbler.com/2.0/artist/beatles/images.rss'); 
$arrFeeds = array(); 
foreach ($doc->getElementsByTagName('item') as $node) { 
    $itemRSS = array ( 
     'title' => $node->getElementsByTagName('title')->item(0)->nodeValue, 
     'desc' => $node->getElementsByTagName('description')->item(0)->nodeValue, 
     'link' => $node->getElementsByTagName('link')->item(0)->nodeValue, 
     'date' => $node->getElementsByTagName('pubDate')->item(0)->nodeValue, 
     'thumbnail' => $node->getElementsByTagName('thumbnail')->item(0)->getAttribute('url') 
    ); 
    array_push($arrFeeds, $itemRSS); 
} 

如果饲料中缺少像thumbnail条目您可能会收到错误Call to a member function getAttribute() on a non-object,因此,虽然我很喜欢@Helder罗巴洛的答案,你应该检查,以确保节点试图用之类的东西getAttribute()之前就存在:

<?php 

header('Content-type: text/plain; charset=utf-8'); 

$doc = new DOMDocument(); 
$doc->load('http://ws.audioscrobbler.com/2.0/artist/beatles/images.rss'); 
$arrFeeds = array(); 
foreach ($doc->getElementsByTagName('item') as $node) { 
    $itemRSS = array (
     'title' => $node->getElementsByTagName('title')->item(0)->nodeValue, 
     'desc' => $node->getElementsByTagName('description')->item(0)->nodeValue, 
     'link' => $node->getElementsByTagName('link')->item(0)->nodeValue, 
     'date' => $node->getElementsByTagName('pubDate')->item(0)->nodeValue 
    ); 

    if(sizeof($node->getElementsByTagName('thumbnail')->item(0)) > 0) 
    { 
     $itemRSS['thumbnail'] = $node->getElementsByTagName('thumbnail')->item(0)->getAttribute('url'); 
    } 
    else 
    { 
     $itemRSS['thumbnail'] = ''; 
    } 

    array_push($arrFeeds, $itemRSS); 
} 


print_r($arrFeeds); 

媒体:内容属性实际上是非常容易得到与简单的XML

if([email protected]$x=simplexml_load_file($feed_url)){ 

} 
else 
{ 
    foreach($x->channel->item as $entry) 
    { 
    $media = $entry->children('http://search.yahoo.com/mrss/')->attributes(); 
    $url = (string) $media['url']; 
    } 
}