CVII. XML parser functions

Introduction

About XML

XML (eXtensible Markup Language)은 웹에서 문서교환을 위해 구조화된 데이터 포멧이다. XML은 The World Wide Web consortium (W3C)에 의해 정의된 기준이고, XML과 그 기술에 관한 정보는 에서 찾아 볼 수 있다.

설치

이 extension은 expat을 사용한다. 이것은 에서 찾아 볼 수 있다. 기본적으로 expat이 포함되어 만들어진 library는 존재하지 않고, 아래의 make rule를 사용하여 만들어 사용할 수 있다.(make파일에 추가.):
libexpat.a: $(OBJS) ar -rc $@ $(OBJS) ranlib $@
expat의 RPM package는 에서 다운 받을 수 있다.

만일 Apache-1.3.7 이후 버젼을 사용한다면 이미 필요한 expat library는 존재함을 명심하라. 단지 PHP를 --with-xml (without any additional path)로 configure 함으로 Apache에 빌드된 expat library를 자동적으로 사용할 수 있다.

UNIX에서는, --with-xml 옵션을 포함하여configure를 실행한다. 컴파일러가 찾을 수 있는 어딘가에 expat library는 설치될 것이다. 만일 Apache 1.3.9 이후 버젼에서 PHP가 모듈로 설치되어 있다면, 아파치에 포함된 expat library를 자동적으로 사용할 수 있다. 만일 expat이 컴파일러가 찾을 수 없는 곳에 설치되어 있다면 CPPFLAGS와 LDFLAGS를 configure 하기전에 설정할 필요가 있다.

PHP를 설치했다. Tada! That should be it (^^).

이 Extension에 관하여.

이 PHP extension 도구들은 James Clark's expat을 지원한다. 이 도구는 XML 문서를 구문대로 분석한다.(하지만 그 유효성은 확이하지 않는다.) 이것은 PHP가 지원하는 세가지 source character encodings을 지원한다. : US-ASCII, ISO-8859-1, UTF-8. 하지만, UTF-16은 지원하지 않는다.

이 extension은 create XML parsers를 생성시키고 다른 XML 이벤트를 위한 handlers를 정의한다. 각각의 XML parser는 또한 조정가능한 parameters를 갖는다.

XML event handlers defined are:

표 1. Supported XML handlers

PHP function to set handler	Event description
xml_set_element_handler()	Element events는 XML parser가 시작과 종료 태그를 만날 때면 발생한다. 이것은 시작태그 handlers와 종료태그 handlers로 구분된다.
xml_set_character_data_handler()	Character data는 거의 XML문서중 태그들 사이의 whitespace를 포함하여 non-markup된 구문이다. XML parser는 whitespace가 유효한지 aplication(혹은 사용자)가 결정하기 전까지는 어느 whitespace도 더하거나 지우지 않음을 명심하라.
xml_set_processing_instruction_handler()	PHP 프로그래머들은 processing instructions (PIs)를 이미 잘 알고 있다. <?php ?>가 processing instruction이고, `php`는 "PI target"이라 불리운다. The handling of these are application-specific, except that all PI targets starting with "XML" are reserved.
xml_set_default_handler()	What goes not to another handler goes to the default handler. You will get things like the XML and document type declarations in the default handler.
xml_set_unparsed_entity_decl_handler()	This handler will be called for declaration of an unparsed (NDATA) entity.
xml_set_notation_decl_handler()	This handler is called for declaration of a notation.
xml_set_external_entity_ref_handler()	This handler is called when the XML parser finds a reference to an external parsed general entity. This can be a reference to a file or URL, for example. See the external entity example for a demonstration.

Case Folding

The element handler functions may get their element names case-folded. Case-folding is defined by the XML standard as "a process applied to a sequence of characters, in which those identified as non-uppercase are replaced by their uppercase equivalents". In other words, when it comes to XML, case-folding simply means uppercasing.

By default, all the element names that are passed to the handler functions are case-folded. This behaviour can be queried and controlled per XML parser with the xml_parser_get_option() and xml_parser_set_option() functions, respectively.

Error Codes

The following constants are defined for XML error codes (as returned by xml_parse()):

XML_ERROR_NONE

XML_ERROR_NO_MEMORY

XML_ERROR_SYNTAX

XML_ERROR_NO_ELEMENTS

XML_ERROR_INVALID_TOKEN

XML_ERROR_UNCLOSED_TOKEN

XML_ERROR_PARTIAL_CHAR

XML_ERROR_TAG_MISMATCH

XML_ERROR_DUPLICATE_ATTRIBUTE

XML_ERROR_JUNK_AFTER_DOC_ELEMENT

XML_ERROR_PARAM_ENTITY_REF

XML_ERROR_UNDEFINED_ENTITY

XML_ERROR_RECURSIVE_ENTITY_REF

XML_ERROR_ASYNC_ENTITY

XML_ERROR_BAD_CHAR_REF

XML_ERROR_BINARY_ENTITY_REF

XML_ERROR_ATTRIBUTE_EXTERNAL_ENTITY_REF

XML_ERROR_MISPLACED_XML_PI

XML_ERROR_UNKNOWN_ENCODING

XML_ERROR_INCORRECT_ENCODING

XML_ERROR_UNCLOSED_CDATA_SECTION

XML_ERROR_EXTERNAL_ENTITY_HANDLING

Character Encoding

PHP's XML extension supports the character set through different character encodings. There are two types of character encodings, source encoding and target encoding. PHP's internal representation of the document is always encoded with UTF-8.

Source encoding is done when an XML document is parsed. Upon creating an XML parser, a source encoding can be specified (this encoding can not be changed later in the XML parser's lifetime). The supported source encodings are ISO-8859-1, US-ASCII and UTF-8. The former two are single-byte encodings, which means that each character is represented by a single byte. UTF-8 can encode characters composed by a variable number of bits (up to 21) in one to four bytes. The default source encoding used by PHP is ISO-8859-1.

Target encoding is done when PHP passes data to XML handler functions. When an XML parser is created, the target encoding is set to the same as the source encoding, but this may be changed at any point. The target encoding will affect character data as well as tag names and processing instruction targets.

If the XML parser encounters characters outside the range that its source encoding is capable of representing, it will return an error.

If PHP encounters characters in the parsed XML document that can not be represented in the chosen target encoding, the problem characters will be "demoted". Currently, this means that such characters are replaced by a question mark.

Some Examples

Here are some example PHP scripts parsing XML documents.

XML Element Structure Example

This first example displays the stucture of the start elements in a document with indentation.
예 1. Show XML Element Structure
$file = "data.xml"; $depth = array(); function startElement($parser, $name, $attrs) { global $depth; for ($i = 0; $i < $depth[$parser]; $i++) { print " "; } print "$name\n"; $depth[$parser]++; } function endElement($parser, $name) { global $depth; $depth[$parser]--; } $xml_parser = xml_parser_create(); xml_set_element_handler($xml_parser, "startElement", "endElement"); if (!($fp = fopen($file, "r"))) { die("could not open XML input"); } while ($data = fread($fp, 4096)) { if (!xml_parse($xml_parser, $data, feof($fp))) { die(sprintf("XML error: %s at line %d", xml_error_string(xml_get_error_code($xml_parser)), xml_get_current_line_number($xml_parser))); } } xml_parser_free($xml_parser);

XML Tag Mapping Example

예 2. Map XML to HTML
This example maps tags in an XML document directly to HTML tags. Elements not found in the "map array" are ignored. Of course, this example will only work with a specific XML document type.
$file = "data.xml"; $map_array = array( "BOLD" => "B", "EMPHASIS" => "I", "LITERAL" => "TT" ); function startElement($parser, $name, $attrs) { global $map_array; if ($htmltag = $map_array[$name]) { print "<$htmltag>"; } } function endElement($parser, $name) { global $map_array; if ($htmltag = $map_array[$name]) { print "</$htmltag>"; } } function characterData($parser, $data) { print $data; } $xml_parser = xml_parser_create(); // use case-folding so we are sure to find the tag in $map_array xml_parser_set_option($xml_parser, XML_OPTION_CASE_FOLDING, true); xml_set_element_handler($xml_parser, "startElement", "endElement"); xml_set_character_data_handler($xml_parser, "characterData"); if (!($fp = fopen($file, "r"))) { die("could not open XML input"); } while ($data = fread($fp, 4096)) { if (!xml_parse($xml_parser, $data, feof($fp))) { die(sprintf("XML error: %s at line %d", xml_error_string(xml_get_error_code($xml_parser)), xml_get_current_line_number($xml_parser))); } } xml_parser_free($xml_parser);

XML External Entity Example

This example highlights XML code. It illustrates how to use an external entity reference handler to include and parse other documents, as well as how PIs can be processed, and a way of determining "trust" for PIs containing code.

XML documents that can be used for this example are found below the example (xmltest.xml and xmltest2.xml.)

예 3. External Entity Example
$file = "xmltest.xml"; function trustedFile($file) { // only trust local files owned by ourselves if (!eregi("^([a-z]+)://", $file) && fileowner($file) == getmyuid()) { return true; } return false; } function startElement($parser, $name, $attribs) { print "<$name"; if (sizeof($attribs)) { while (list($k, $v) = each($attribs)) { print " $k=\"$v\""; } } print ">"; } function endElement($parser, $name) { print "</$name>"; } function characterData($parser, $data) { print "$data"; } function PIHandler($parser, $target, $data) { switch (strtolower($target)) { case "php": global $parser_file; // If the parsed document is "trusted", we say it is safe // to execute PHP code inside it. If not, display the code // instead. if (trustedFile($parser_file[$parser])) { eval($data); } else { printf("Untrusted PHP code: %s", htmlspecialchars($data)); } break; } } function defaultHandler($parser, $data) { if (substr($data, 0, 1) == "&" && substr($data, -1, 1) == ";") { printf('%s', htmlspecialchars($data)); } else { printf('%s', htmlspecialchars($data)); } } function externalEntityRefHandler($parser, $openEntityNames, $base, $systemId, $publicId) { if ($systemId) { if (!list($parser, $fp) = new_xml_parser($systemId)) { printf("Could not open entity %s at %s\n", $openEntityNames, $systemId); return false; } while ($data = fread($fp, 4096)) { if (!xml_parse($parser, $data, feof($fp))) { printf("XML error: %s at line %d while parsing entity %s\n", xml_error_string(xml_get_error_code($parser)), xml_get_current_line_number($parser), $openEntityNames); xml_parser_free($parser); return false; } } xml_parser_free($parser); return true; } return false; } function new_xml_parser($file) { global $parser_file; $xml_parser = xml_parser_create(); xml_parser_set_option($xml_parser, XML_OPTION_CASE_FOLDING, 1); xml_set_element_handler($xml_parser, "startElement", "endElement"); xml_set_character_data_handler($xml_parser, "characterData"); xml_set_processing_instruction_handler($xml_parser, "PIHandler"); xml_set_default_handler($xml_parser, "defaultHandler"); xml_set_external_entity_ref_handler($xml_parser, "externalEntityRefHandler"); if (!($fp = @fopen($file, "r"))) { return false; } if (!is_array($parser_file)) { settype($parser_file, "array"); } $parser_file[$xml_parser] = $file; return array($xml_parser, $fp); } if (!(list($xml_parser, $fp) = new_xml_parser($file))) { die("could not open XML input"); } print "<pre>"; while ($data = fread($fp, 4096)) { if (!xml_parse($xml_parser, $data, feof($fp))) { die(sprintf("XML error: %s at line %d\n", xml_error_string(xml_get_error_code($xml_parser)), xml_get_current_line_number($xml_parser))); } } print "</pre>"; print "parse complete\n"; xml_parser_free($xml_parser); ?>

예 4. xmltest.xml
<?xml version='1.0'?> <!DOCTYPE chapter SYSTEM "/just/a/test.dtd" [ <!ENTITY plainEntity "FOO entity"> <!ENTITY systemEntity SYSTEM "xmltest2.xml"> ]> <chapter> <TITLE>Title &plainEntity;</TITLE> <para> <informaltable> <tgroup cols="3"> <tbody> <row><entry>a1</entry><entry morerows="1">b1</entry><entry>c1</entry></row> <row><entry>a2</entry><entry>c2</entry></row> <row><entry>a3</entry><entry>b3</entry><entry>c3</entry></row> </tbody> </tgroup> </informaltable> </para> &systemEntity; <sect1 id="about"> <title>About this Document</title> <para>  <?php print 'Hi! This is PHP version '.phpversion(); ?> </para> </sect1> </chapter>

This file is included from xmltest.xml:
예 5. xmltest2.xml
<?xml version="1.0"?> <!DOCTYPE foo [ <!ENTITY testEnt "test entity"> ]> <foo> <element attrib="value"/> &testEnt; <?php print "This is some more PHP code being executed."; ?> </foo>

차례
utf8_decode -- Converts a string with ISO-8859-1 characters encoded with UTF-8 to single-byte ISO-8859-1.
utf8_encode -- encodes an ISO-8859-1 string to UTF-8
xml_error_string -- get XML parser error string
xml_get_current_byte_index -- get current byte index for an XML parser
xml_get_current_column_number -- Get current column number for an XML parser
xml_get_current_line_number -- get current line number for an XML parser
xml_get_error_code -- get XML parser error code
xml_parse_into_struct -- Parse XML data into an array structure
xml_parse -- start parsing an XML document
xml_parser_create_ns -- Create an XML parser
xml_parser_create -- create an XML parser
xml_parser_free -- Free an XML parser
xml_parser_get_option -- get options from an XML parser
xml_parser_set_option -- set options in an XML parser
xml_set_character_data_handler -- set up character data handler
xml_set_default_handler -- set up default handler
xml_set_element_handler -- set up start and end element handlers
xml_set_end_namespace_decl_handler -- Set up character data handler
xml_set_external_entity_ref_handler -- set up external entity reference handler
xml_set_notation_decl_handler -- set up notation declaration handler
xml_set_object -- Use XML Parser within an object
xml_set_processing_instruction_handler -- Set up processing instruction (PI) handler
xml_set_start_namespace_decl_handler -- Set up character data handler
xml_set_unparsed_entity_decl_handler -- Set up unparsed entity declaration handler


	downloads \| documentation \| faq \| getting help \| mailing lists \| \| php.net sites \| links \| my php.net

search for in the